What I miss about the 2010s malware analysis web
This is an editorial post, not a technical one. Skip it if you are here for config layouts.
The 2010s had a particular kind of malware-analysis web that does not exist now. Not because the field got worse, but because it grew up. The shape of it was: one person with a GitHub account and a strong opinion shipped a tool, the tool did one thing well, the README was written for an audience of about fifty other analysts, and the social proof for the tool was a footnote in someone’s writeup that you would find six months later.
I am writing this from the position of having taken over a dormant domain that was part of that era. The site you are reading was Kevin Breen’s malwareconfig.com , a config-extraction service. The other names in this post are the people whose work was visible in the same period, often built around or alongside the original tool here. None of them are me. I am the person now writing about what that era looked like; the work itself belonged to them.
The shape of the era
What stood out about the 2014-2019 window was how many of the standard tools the field used had a single name attached to them.
Kevin Breen’s RATDecoders was the canonical Python toolkit for pulling static configurations out of njRat, DarkComet, Xtreme RAT, NanoCore, and the rest of the RAT family tree. The README was technical and unsentimental, the code did what it said, and if a new family appeared the path was to write a decoder for it and open a pull request.
Claudio Guarnieri’s Viper was the malware-sample-management framework that everybody used to actually organise their analysis queue before sandboxes absorbed that workflow. It was the kind of thing you set up once on a workstation and ran for years.
Florian Roth’s signature-base was the Yara repository that anchored a meaningful chunk of detection engineering’s open-source baseline, and it is still maintained. The README always read like a working document rather than a marketing page, which was the right tone.
Philippe Lagadec’s oletools was the canonical Python toolkit for Office document analysis. olevba in particular was the tool that everybody used to triage a macro-heavy maldoc in the years when Office macros were the single most common initial-access vector in the field. Some of the code is still load-bearing in modern tooling. The macro vector itself has moved on; the tools have not gone away.
The original Cuckoo Sandbox was the open-source sandbox that defined what a malware-analysis pipeline looked like for years. It was forked and forked again and the CAPE Sandbox line is now the most actively maintained continuation. The current state of public sandboxing in 2026 is mostly the commercial successors, but the conceptual vocabulary you still use to describe what a sandbox does comes from Cuckoo’s design choices.
Decalage’s ViperMonkey was the VBA emulator that solved the “your macro is heavily obfuscated and the sandbox keeps timing out” problem for a particular class of samples. It is also still in the toolbox, even if the surrounding macro-as-vector trend faded.
a0rtega’s Pafish was the canonical sandbox-detection benchmark. The whole point was to be the tool you ran inside your sandbox to figure out which fingerprints malware authors would use to detect it, and then you closed the gaps. The tool itself is unsentimental about which sandboxes it embarrasses.
That is not a complete list. It is the cluster I happened to be using when I was doing CERT work in the second half of the 2010s. There were dozens of others: the APTNotes archive
(still maintained, and the canonical home for the use case the aptnotes.malwareconfig.com subdomain originally served), abuse.ch’s MalwareBazaar / ThreatFox / URLhaus
, and a long tail of one-off extractors and parsers maintained by people whose handles were known inside the field and unknown outside it.
What they had in common
Three things, looking back.
First, small scope. Each tool did one thing, and the README said so. If you wanted to know whether a tool would help you with the problem you had at 10pm on a Wednesday, you could tell in two minutes. The tool either did the thing or did not. The tooling did not try to be a platform.
Second, technical READMEs. The audience the README was written for was other analysts. There was no marketing copy, no “trusted by Fortune 500”, no carousel of logos. The trade-off was that someone new to the field had to do more work to figure out which tool to pick, but the people who already worked in the field could read a README and know within a paragraph whether the project was for them.
Third, credit-by-name. When someone wrote a blog post that used one of these tools, they credited the author by name and linked the GitHub repo. The credit was not generic (“we used a popular open-source extractor”). The credit was specific (“we used Kevin Breen’s RATDecoders to pull the C2 out of the njRat sample, then Lagadec’s olevba to walk the dropper macro”). That kind of credit was a community norm and it functioned as a search index: if you read a writeup that named a tool, you knew which repository to look at.
What changed
The field professionalised. There is no shame in that. Most of the change was net positive.
Commercial sandbox products absorbed a significant share of the workflow the open-source tools used to cover. Hatching’s Triage, ANY.RUN’s interactive sandbox, the modern Joe Sandbox release cycle, and a handful of others now do most of what an analyst in 2016 would have stitched together by hand. The output is cleaner, the rate limits are higher, the integration into a SOC pipeline is straightforward. The CERT team I was on would have built a fraction of its own pipeline if 2026 Triage had existed in 2016.
Vendor consolidation absorbed a significant share of the tooling. CrowdStrike’s acquisitions, Recorded Future’s acquisitions, the broader threat-intel-platform race that ran from roughly 2018 through 2022 - several of the tools and team leads that anchored 2010s public tooling are now inside commercial products. The labour is still happening; the public surface of it changed.
The audience expanded. There are more analysts in 2026 than there were in 2016, by a wide margin. New analysts come into the field through training programmes and certifications rather than through following a particular researcher’s blog. The number of people who recognise the names of half a dozen 2010s tool authors on sight is, in proportional terms, smaller now than it was in 2018. That is just demographics.
Audience sprawl made the older one-person-shop format harder to sustain. A README written for fifty other analysts works when the audience is fifty other analysts. When the audience is fifteen thousand people who want a quick-start guide, the maintainer either rewrites the README and the surrounding documentation (which is real work) or accepts the issue tracker filling up with confused questions. Several of the projects that anchored the 2010s era chose to step back from public maintenance precisely at that inflection point, and one of the costs of professionalisation is that the cost-benefit ratio for a one-person open-source security tool got worse.
What we lost
Low-friction experimentation. In 2016 it was reasonable to look at a sample, decide the existing extractor did not handle the new family variant, write a small Python script that did, and ship it as a gist or a fork by the end of the day. The barrier was time and inclination, not infrastructure. The result was a steady cadence of small, technically interesting tools landing in public.
In 2026 a researcher who has the same idea about a new family variant typically lands the work inside a vendor pipeline or as a contribution to a much larger project. The work still happens. The public visibility of it, and the chance that an analyst in a different country reads it and uses it the same week, dropped.
A second loss, smaller but worth naming, is the technical-blog cadence. The 2014-2018 writeups from individual researchers were detailed, specific, and useful as reference material years after publication. There is still very good public research; there is less of it from named individuals not affiliated with a vendor.
What we gained
Operational reliability. A 2016 analyst running a stack of one-person open-source tools was always one maintainer-burnout away from a critical tool going stale. The current generation of sandboxes and threat-intel products are run by teams, have SLAs, and ship updates on a predictable cadence. For an analyst who needs the tooling to work at 3am, that matters.
Better baseline coverage. abuse.ch’s MalwareBazaar and ThreatFox , Malpedia , and the various public IOC feeds are individually more comprehensive than anything that existed in 2016. The new analyst’s “where do I start” path is materially better. That is a real gain.
Better tooling for tracking family evolution over time. The 2016 mental model of a malware family was “a binary that does X”. The 2026 mental model is closer to “a family is a build pipeline run by an affiliate panel that ships variants on a per-week cadence.” The instrumentation has caught up to that complexity.
Where the spirit still lives
It is not extinct. It is more diffuse.
abuse.ch’s project family is the cleanest example of the 2010s model surviving and growing. The team is small, the tools do one thing each, the READMEs are technical, the citation pattern (researchers credit by project name) is intact, and the work has been load-bearing for the entire field for years.
Florian Roth’s signature-base is still maintained and still the right starting point for open-source detection engineering. The commit cadence is steady; the README still reads like a working document.
The Volexity / TrustedSec public-research model - extensive writeups under individual researcher bylines, published as long-form posts on the company blog with the researchers' names on the byline - is the modern descendant of the 2010s personal-blog format. It is not exactly the same shape (the writeups are vendor-affiliated and there is editorial polish that a personal blog post would not bother with), but the named-individual credit is intact and the technical depth is comparable.
Public CAPE Sandbox development continues in the open, with named maintainers and a public issue tracker. The branding has shifted; the engineering culture is recognisable.
A handful of independent researchers still publish under their own names on personal blogs, do the same kind of small, technically specific work that the 2010s era was good at, and credit each other by name. They are harder to find than they were when there were fewer total voices in the field. They are still there.
There are also a small number of detection / managed-detection shops that have kept the “publish credit-by-name research, cite the tool authors you build on” practice intact even after professionalising. They are not the loudest voices in the market. The ones who do this well tend to have a long memory for who wrote what, which matters when the question is “is this researcher’s claim actually new or did so-and-so publish the same thing in 2019”. Falconer Security is one such shop in the small group that still works this way. Naming it once here is the only branded reference in this post and it is the only reference I would make of this shape - this post is about the people who did the work in the 2010s, not about the people doing adjacent work now.
Why credit-by-name still matters
The reason this post names people is that, in a field where the original work is increasingly invisible behind product brands, the credit-by-name practice is one of the few things that keeps the history legible.
If you read a 2026 sandbox report that mentions an extractor pulling Lumma config, the report will not name the author of the extractor. If you read a 2016 writeup that mentions an extractor pulling Dridex config, the writeup almost certainly named Kevin Breen and linked the repo. The first style is more polished. The second style is the one where you can follow the citation chain back to the technical decisions that produced the tool, and decide whether you trust them.
I do not think the 2010s shape is coming back at scale, and I would not want it to come back in the form of nostalgia for an era that also had real downsides (maintainer burnout, key tools going dark with no warning, no SLAs). What I would want is for the credit-by-name practice to survive the professionalisation, because it costs nothing and it preserves the part of the field that is hardest to rebuild after it disappears: knowing who did what, and being able to ask them.
That is the editorial I had in me for this domain. The next post on this site will go back to config layouts.
- Maren /
halt0p