tecosystems

Who’s Writing Open Source Code?

Share via Twitter Share via Facebook Share via Linkedin Share via Reddit

As AI steadily reshapes the software industry around it as open source once did before it, it’s useful to try and understand where those two forces intersect. Open source has, as it has with virtually every other industry software category, had an enormous impact on AI. AI offerings are built on vast foundations of open source, and an emerging new set of highly capable models have been released in a manner reminiscent of open source, if not open source the way we’ve traditionally understood it. This has led to much recent vendor jockeying, from unusually coordinated defenses of open weight models to defensive critiques. But that’s a topic to be tabled for the moment, at least until some of the dust has settled.

Instead, as the impact of open source on AI makes headlines, it’s useful to examine the reverse: what is the impact of AI on open source? This is a common question in the industry, both for individual developers and maintainers, those responsible for project governance as well as vendors that sponsor open source work in some capacity. In the wake of examples like Bun, in which the project has transitioned from primarily being written by humans to one authored by machines, is that an outlier or the new norm? Put more simply: are open source projects more broadly still written by humans, or have the robots taken over?

Setting aside for the moment subjective feedback from project maintainers about the impact of AI, which in general is grim, what can we understand objectively and analytically about how AI is or isn’t contributing to open source code? To try and answer that, 15 projects were selected and analyzed to see what evidence they can provide for the growth or lackthereof of AI-authored code.

The Caveats

Sadly, the truth is that what we can say objectively, at any scale beyond a single project, is limited. We cannot, for example, account for code written by a machine but passed off by a human as their own work. Bot-delivered human code, likewise, can be difficult to identify and parse in some cases. Even bots themselves are limited to the hardcoded regexes that detect them. We cannot state with any degree of precision, then, what the exact percentage of code that was written by a human versus that of a machine.

We can, however, establish something of a floor from a contribution standpoint.

In simple terms, what can’t be stated is how much actual AI code is in these projects. What can be captured, however, is the number of commits that carry a machine-readable AI marker: a co-author trailer naming an AI tool, or authorship by an autonomous agent. Even those, however, could have plausibly different explanations.

  1. New Instrumentation: it could simply be a function of the sudden rise of Claude Code appending “Co-authored-by: Claude” or Copilot’s agent mode doing the same thing.
  2. New Tools: Copilot, for example, has been in market and assisting in development since 2021, but only in the last 18 months have agentic CLI tools really emerged as autonomous or semi-autonomous collaborators. In other words, growth is less a function of reporting and more a reflection of exploding usage of a new class of coding assistance tools.

In the end, however, both are likely to be true and there is no realistic way of apportioning them by way of this dataset. All that is available are coarse markers of overall usage. We can also, however, continue to take these coarse snapshops to monitor the slope of that minimum floor to see if usage is increasing, decreasing or remaining static.

The Study Cohort

To begin with, of course, a sample is needed. To determine project size, this analysis relies on the size of the contributor base. It’s an imperfect metric for that purpose, of course, but they all are. Next, we looked at criticality – i.e. how important or used a given project was – as measured by a variety of metrics. For medium and large projects, it was measured via deps.dev dependencies, with a fallback to GitHub mentions. Small projects, for their part, were less amenable to this approach and were instead measured via Debian Popcon popularity. Ultimately, while the sample wanted to reflect projects of differing sizes, the intent was to examine projects that were being used and relied upon.

To add one further linking thread, or remove one potential confounding variable, this analysis prioritized having projects written in a common language in all three cohorts. Given that a number of the most critical small projects were written in C, that was prioritized and that’s a bias worth being aware of here. Future runs will prioritize other more modern languages to assess whether there are language distinctions in AI uptake.

The only other thing to note here is that the data used in this analysis was obtained via the GitHub API rather than the GitHub Archive due to significant data loss issues with the latter beginning in 2025. This means, among other things, that sample projects had to be hosted on GitHub for the entire period of time under scrutiny – projects hosted on mirrors (e.g. Postgres) or which experienced migrations (e.g. xz) were thus excluded. The GitHub API has the advantage of being able to return 2026 data, but it is subject to survivorship bias in a way that the Archive is not. Force-pushed commits, squash merges and more are all not captured by the API as they were with the Archive, so bear that in mind as well.

With all of those caveats out of the way, above are the cohorts – five sample projects per category.

The Results

All Projects

The single biggest takeaway from examining AI-created code amongst 15 critical infrastructure projects is that there is not provably a lot of it. Across ~24K commits made in the three categories over the first half of 2026, for example, the number authored by a machine measured in the dozens. Less than 1% of the surveyed code was inarguably written by machines.

It’s worth stressing again that these numbers can only be off in one direction – at the expense of AI-authored code, and that it’s almost certain that visibility is lagging usage. But for those worrying that open source codebases were overrun by AI slop, that has not happened yet – though from reports that’s only because it is not being committed by maintainers.

Usage by Project Size

The next obvious question was whether AI uptake varied by project size. While acknowledging that with such a small sample of actual commits it would be impossible to come to a confident conclusion in either direction, the data that we do have suggests there’s little distinction between the size cohorts in their affinity for AI-code.

The Small bar’s sample size – less than 1K commits – means that it’s statistically meaningless. But while the Medium product category seems to be more open to AI-authored code, the Y axis is honest about the scale at work and beyond that, over 80% of Medium’s commits come from a single project (esp-idf).

Usage by Project

To understand in more detail which projects in the cohort were committing AI-generated code and in what degree relative to one another, here is a chart of the commits by both project size and the projects themselves.

Usage as measured by project, presumably exacerbated by the small sample, was uneven. Two projects count for 73% of the total commits, almost half have no documented AI commits full stop.

AI Agents vs Commits w/ AI Assistance

Of some interest was the relative count of commits from autonomous AI agents versus commits publicly acknowledging coding assistance, which at present is tilted in favor of the latter.

For all the talk of fleets of autonomous agents marshaled by an emerging crop of next generation IDEs for architects, whatever contributions they are making are not being recorded explicitly in any volume as yet.

Takeaways

Due to the data limitations, the following conclusions should be considered speculative at best. But based on the present day snapshot, the following seem the most significant takeaways.

  • Floor: the first half of 2026 is likely to be the lowest in recorded, visible AI-contributions for the foreseeable future. It’s lower in part because unreported usage is going to be more common than reported, as discussed above, but also because if one assumes the AI genies will not be put back into the bottle, there is only one direction these charts will move in.
  • Signal: it’s possible that the transition from no usage to usage in 2026 is a genuine reflection of new and growing usage of coding assistants and acceptance of their output into open source projects. It is more likely to be a function of the fact that reporting of AI generated code is a subset and more likely a small fraction of actual code contributed thanks to the emergence of emitting trailers. In either case, however, the signal is what’s important – the cause is more of an academic concern.
  • Distribution: As William Gibson says, the future is here, it’s just unevenly distributed. As mentioned above, the measured commits here aren dominated by two out of the fifteen projects, while around half aren’t using it at all. Adoption is not likely to be systemic, in other words, it will be dynamic and highly asymmetric – driven by a variety of localized factors from policy, tooling to a particularly enthusiastic maintainer.
  • Humans: it presumably goes without saying given the extremely limited amount of AI code detected, but there is no measurable impact to human code contributions as yet. If usage expands as expected, that might change, but at present across these projects there is, on average, no impact to human contribution rates.
  • Transparency: it should not be controversial to state that coding assistance had a step change in functional capabilities as of November 2025, and that usage exploded following that. This dataset only reflects the early stages of that expansion, but even so its inherent limitations point to a tension that open source communities are already and will continue to grapple with: transparency. Historically, the effort required to generate code meant that submitting code to an open source project, for whatever purpose, had a cost. As that cost has gone down, the barriers to contributing have fallen, but the risks to the project – from provenance to security – have not, and more importantly nor has the developer incentives to make contributions. This will inevitably lead to scenarios in which users will contribute code they did not write without disclosing that fact, which poses challenges for researchers and the project communities alike. Transparency into code origin, in other words, should only become more important from a cultural and etiquette standpoint.