A few months ago I decided to build an idea I’d been considering for a while — the ability to scan an entire project for license information (e.g. LICENSE.md files) and report back any possible conflicts with a company’s organizational policy (e.g. a GPL-3.0 dependency prohibited by company policy).
To be clear, the idea I had was not necessarily a novel one. In fact, years ago, when I had been working at a previous startup, I had been tasked to create this exact kind of license report, and had found a bunch of disparate tools — some free, some paid. The challenge was that the free tools weren’t that good and were usually ecosystem-dependent, and the paid tools were outside our budget. So, my solution at the time was a mix of using a bunch of the free tools, compiling and normalizing the reports they created into one big report, and then manually fixing and adjusting things until it all seemed cohesive. I made a mental note that one day I’d build a great open source alternative—and potentially build a business around it that undercut the paid tools.
Within a month of coding and working with Claude, I had done it: https://licensedetector.com/ and https://github.com/licensedetector/cli/ both exist, and the tool is better than anything out there that I’ve encountered, including the 2,600-star ScanCode Toolkit library as shown below1:
| Metric | License Detector | askalono | ScanCode | licensee |
|---|---|---|---|---|
| SPDX-ID accuracy | 99.0% | 96.9% | 92.8% | 35.1% |
| Throughput (files/sec) | 161/sec | 9/sec | 16/sec | 2/sec |
| Cold start | 235 ms | 108 ms | 4.6 s | 324 ms |
| Adversarial survival | 98% | 91% | 86% | 21% |
| …on truncated text | 89% | 39% | 32% | 0% |
After building the tool, however, I found myself a bit unsatisfied — could I have put that same amount of effort into contributing to ScanCode? Thinking it through, I believe the answer is no — and it makes me worry for the state of open source.
To explain, there are a few challenges with using existing open source libraries, both in using the libraries and in contributing back to them:
- Using the libraries:
- Having to deal with legacy code that can’t be removed / changed because of choices made years ago
- Having to deal with the language and architecture chosen by the maintainers, even though they may not ergonomically fit your needs
- Contributing back:
- Having to “convince” the maintainers that your feature / use cases are warranted
- Overwhelming the maintainers with large refactors that might improve the tool but are difficult to review
- Messing up someone’s workflow:

The allure of a greenfield project is just too much to pass up — you can build the tool that you want, and not spend time convincing people that your use case is worth supporting.
That all being said, where does this leave open source? If someone can spend a month and build an objectively better tool than the current open source leader, why would anyone contribute to the existing project? There are a few obvious reasons why, and they are still important:
- Working with others: Open source is not always about the quality of the code; it’s about working with other engineers in the open. Showing you can work well with others is a great way to show employers you have good communication skills, and will work well in a team.
- Building towards something unified: It’s easy as engineers for us to say “all these existing libraries suck, I’m going to build my own”, but you run into the problem where rather than making it easier for others to see which library is best, you add to the glut. If a group of people worked on something together, though, it could have a chance to really make a change in the ecosystem.

Still, something in open source will need to change — or we as engineers will have to treat open source code differently. With the proliferation of AI tooling and cheap compute, it will be easy for engineers to create open source libraries that beat the best ones out there but receive only a handful of stars. When an engineer comes along and wants to look for the best tool out there, how will they decide between the 2,000-star repo that is actively being worked on but is slow and old, and the 300 repos that people have created that are actually great, but were created by one person and never maintained? Or will that engineer simply say “I guess I’ll build my own tool from scratch”, and add to the glut?
The clearest parallel to this is the explosion of writing that happened after the printing press became widely available. There’s great literature on how people of that time dealt with this2, but in a nutshell, readers relied on trusted institutions and individuals for book reviews and summaries. Borrowing books from libraries or friends also reduced the financial cost of choosing a bad book.
Can we apply these same principles to code today? To be clear, we already do this to some extent — GitHub repos with “Awesome X” names collect the best libraries around topic X, AI tools summarize content and surface useful projects through search, and YouTubers and other content creators recommend new libraries and frameworks. But this still feels a bit unsatisfying — the pace at which we are creating new libraries and content vastly exceeds the capacity of both AI and humans to consume it.
But where does this leave us then? In my mind, I have to believe it creates a divide between libraries / tools / frameworks that are already well known and everything else. For the libraries and tools that are already heavily used and well known, they will continue to get open source support, even if a greenfield project is better / faster / has more functionality — distribution just wins. For the tools that come afterwards, unless you can figure out how to market your tool, the best you’ll be able to say is that your code will become fodder for AI models that will allow the next developer to come along and build their own tool with the help of AI — perhaps not as enticing as one might hope.
Footnotes:
- Don’t believe me? I included the benchmark tool for you to run yourself: https://github.com/licensedetector/cli/tree/main/bench ↩︎
- See: Too Much to Know: Managing Scholarly Information before the Modern Age by Ann M. Blair ↩︎