The attribution problem: why nation-state claims are probabilities
By Elias Lankinen11 min read
I have thorough research across primary sources and multiple cases. Now I'll write the post.
The night the Olympics went dark
At 8 p.m. on 9 February 2018, as the opening ceremony of the PyeongChang Winter Olympics filled a stadium in the South Korean mountains, the technology underneath the spectacle quietly collapsed. Wi-Fi died. The official ticketing app stopped working, stranding spectators at the gates. Broadcast drones were grounded, and monitors in the press centre went black. Backstage, engineers watched servers wipe themselves. A piece of malware later named Olympic Destroyer was tearing through the Games' IT backbone, and the people paid to defend it spent the night wrestling systems back online rather than watching the fireworks. Then came the strange part. When Kaspersky's researchers pulled the malware apart, they found a fingerprint pointing squarely at North Korea, a distinctive fragment of code that matched tools used by the Lazarus Group, Pyongyang's state hackers. It was almost too convenient. On closer inspection, that fingerprint had been forged: someone had hand-crafted a section of the file's metadata to imitate North Korean code that had nothing to do with the rest of the program. The clue was a plant, deliberately laid to send investigators chasing the wrong country.
This is the attribution problem in a single incident. Two years later the United States charged six officers of Russia's GRU military intelligence agency with the attack, part of the Sandworm unit. But that conclusion was never a fact recovered from a hard drive. It was a judgment, built from many pieces of evidence, some of them deliberately poisoned, and carrying a stated probability of being wrong. When a government says a nation-state was behind a cyberattack, it is not reading a verdict off a screen. It is placing a bet, and telling you the odds.
What attribution actually means
The word hides three separate questions that get answered with wildly different confidence. Security researchers sometimes describe attribution as a set of nested dolls. The innermost question is technical: which machine or piece of malware did this? You can often answer that with near-certainty from logs and forensic samples. The middle question is: who was operating that machine — which human, or which named hacking crew? That is much harder. The outer question, the one that ends up in headlines and sanctions, is: on whose behalf? Which government directed or paid for the operation? That is the hardest of all, because a keyboard in Shanghai or St Petersburg does not come with a chain of command attached. A common misconception collapses these into one. People assume that because the internet is traceable, attribution is basically a lookup, or, taking the opposite view, that because attackers route through hijacked computers in a dozen countries, attribution is impossible and any government accusation is theatre. Both are wrong. The truth is that attribution is a probabilistic inference that combines forensics with everything else an intelligence service knows, and its confidence ranges from near-worthless to overwhelming depending on the case. The evidence falls into three rough layers. Technical indicators are the artefacts: IP addresses, malware samples, the command-and-control servers the attackers phone home to, reused encryption keys. Behavioural indicators, often called TTPs (tactics, techniques and procedures), are how a group works: the specific tools it favours, the hours it keeps, the sequence of moves it makes once inside a network. Contextual indicators are the strategic fit: who benefits, what languages appear in the code, which targets align with a state's known intelligence priorities. No single layer is decisive. The confidence comes from watching them converge.
Why certainty is off the table
The Olympic Destroyer forgery illustrates the central obstacle: the most objective-looking evidence is the easiest to fake. An IP address can be a rented server in a third country. A malware sample can be copied, recompiled, or salted with another group's code. Time-zone stamps and keyboard-language settings can be set to anything. This is the false flag problem, and it is not hypothetical. Kaspersky's Igor Soumenkov, who unpicked the Olympic Destroyer forgery, called it one of the best he had ever seen precisely because the fake North Korean fingerprint was embedded in a place analysts trust. CIA documents released via WikiLeaks in 2017, the "Vault 7" leak, revealed an internal toolkit called Marble that could obfuscate the agency's own malware and even insert misleading foreign-language strings. If one major intelligence service builds that capability, prudent analysts assume others have too. The upshot is that any indicator an attacker could plausibly have controlled must be discounted, and the strongest evidence is usually the mistake the attacker did not intend to leave. That is why good attribution leans hard on behaviour and on operational slip-ups rather than on the artefacts an adversary chooses to show you. Malware can be swapped out overnight; the habits of a team of human operators, working under deadlines, are far stickier and far harder to fake convincingly.
The Q Model: attribution as judgment, not forensics
The most influential attempt to make sense of this came in 2015, when Thomas Rid and Ben Buchanan of King's College London published Attributing Cyber Attacks in the Journal of Strategic Studies. Their core line, quietly radical at the time, was that "attribution is what states make of it." Attribution, they argued, is not a technical fact waiting to be uncovered but a process that produces a judgment, shaped at every stage by human choices. Their framework, the Q Model, splits the work into three levels. At the tactical level, attribution "is an art as well as a science": forensic investigators piece together malware and logs, but interpreting them requires experience and inference. At the operational level, it "is a nuanced process, not a black-and-white problem": analysts weigh incomplete, sometimes contradictory evidence and assign confidence. At the strategic level, attribution "is a function of what is at stake politically": whether and how a government announces a conclusion depends on what it wants to achieve. Two implications follow. First, attribution improves with time and effort rather than arriving all at once, which is why an initial "we don't know" can honestly become "we are highly confident" months later without anyone having lied. Second, as Rid and Buchanan put it, communicating attribution is itself part of attributing. The decision to name a culprit, and how much evidence to reveal, is inseparable from the act of attribution. That reframing, from forensics to judgment, is why nation-state claims are best read as probabilities.
The language of maybe
If attribution is a probability, the honest thing to do is state the odds, and the intelligence world has spent sixty years learning how badly that goes when you leave it vague. In 1964 the CIA analyst Sherman Kent wrote a now-famous essay, Words of Estimative Probability, after discovering that colleagues reading the same estimate, that something was a "serious possibility", interpreted it as anywhere from a 20 percent to an 80 percent chance. A phrase everyone thought they understood was smuggling a fourfold disagreement. Kent's fix, standardising verbal probabilities against rough numeric ranges, eventually hardened into formal doctrine. The US intelligence community's Directive 203, first issued in 2007 and last amended in January 2022, sets the analytic standards every agency must follow. It draws a sharp distinction between two things that ordinary language blurs. Likelihood is how probable the event is: "almost certainly," "likely," "roughly even chance," and so on, each mapped to a percentage band. Confidence is how good your evidence is, expressed as high, moderate or low. The directive is explicit that a high confidence judgment "is not a fact or a certainty" and "still carries a risk of being wrong," and that moderate confidence means the information is credibly sourced and plausible but not corroborated well enough to say more. Crucially, ICD 203 forbids analysts from combining a confidence level and a likelihood term in the same sentence, because doing so muddles what exactly is uncertain: the world, or your read on it. This is why the phrasing of attribution statements is not bureaucratic hedging. When Britain's National Cyber Security Centre said the Russian military was "almost certainly" responsible for an attack, "almost certainly" was a specific probability band, chosen deliberately over "certainly." Reading these words as weasel language misses that they are the most precise part of the claim.
When the evidence gets overwhelming
None of this means attribution is hopeless. Probabilities can climb very high, and they do so when attackers make human mistakes that no amount of technical misdirection can undo. Take the 2016 breach of the US Democratic National Committee. A persona calling itself Guccifer 2.0 claimed to be a lone Romanian hacker with no state ties. But as The Daily Beast reported, on at least one occasion the operator forgot to switch on the VPN that hid his location, leaving a single honest connection that traced back to a specific GRU officer working out of the agency's headquarters on Grizodubovoy Street in Moscow. In July 2018 the Special Counsel's office indicted twelve GRU officers and named the Guccifer 2.0 persona as their creation. A slip of tradecraft turned a deniable persona into a named intelligence unit. The 2014 Sony Pictures hack shows the same pattern under fire. When the FBI attributed the destruction of Sony's network to North Korea, skeptics called the evidence flimsy, and some argued for a disgruntled insider or Russian-speaking authors. But the malware overlapped with tools used in earlier attacks on South Korean banks, and investigators found that the attackers had occasionally connected directly from IP addresses used exclusively by North Korea, apparently through sloppy operational security. FBI Director James Comey said he had "very high confidence" in the conclusion, a phrase, not a proof.
The high-water mark of confident attribution came a year earlier. In February 2013 the private firm Mandiant published a report on a group it called APT1, tracing a seven-year campaign that stole hundreds of terabytes from at least 141 organisations to a single unit of China's People's Liberation Army, Unit 61398, and even to a specific twelve-storey building on Datong Road in Shanghai. Mandiant did it by watching the operators' habits over years and catching the moments they logged into their own social-media accounts from the same infrastructure they used to attack victims. It was the first time a company had publicly pinned a sustained espionage campaign on a named military unit, and it effectively created the modern threat-intelligence industry.
Attribution is a political act
Notice how much of that evidence never becomes public. Governments sit on intercepts, informants and implants they cannot reveal without burning the source. So a state's attribution claim asks you to trust not only the analysis but the decision about how much to show, and that decision is political. The clearest example is NotPetya, the malware that spread from Ukrainian accounting software in June 2017 and caused more than $10 billion in global damage, the most costly cyberattack in history. Public attribution did not come for eight months, and when it did in February 2018 it arrived as a coordinated announcement by the United States, the United Kingdom, Australia, Canada, Denmark, Estonia, Lithuania, New Zealand and Norway, all blaming the Russian military on the same day. The delay was not analytical hesitation so much as diplomatic choreography. Naming a nation-state is a foreign-policy act with consequences, so the timing, the wording and the roster of co-signers are all chosen. Russia dismissed the charge as a "Russophobic campaign" lacking evidence, which is the predictable other half of the ritual.
Indictments are a related instrument. When the US Justice Department charged five PLA officers in 2014 and later Russian and North Korean operatives, it knew none of them would ever see a US courtroom. The point was to put the evidence on a legal record, signal to adversaries that they had been identified, and establish attribution in a forum with a high standard of proof. Attribution here is less about punishment than about turning a probability into an official, on-the-record accusation.
That some of these actors are named by private companies, chasing headlines and clients, adds another wrinkle. Firms like Mandiant, CrowdStrike and Kaspersky now issue attributions that shape public opinion before any government speaks, and their incentives, marketing, geography, the flags of their home countries, are not neutral. A Russian firm exposing an American operation and an American firm exposing a Russian one are both doing real forensic work and both operating inside a political frame.
What to watch
The honest way to read a nation-state attribution, then, is the way you would read a weather forecast or a medical test result: as a calibrated probability from a source with a track record, not as a photograph of the truth. The right follow-up questions are never "is this certain?" but "how confident, on what evidence, and revealed by whom?" Two trends will make this harder. Generative AI lowers the cost of forging the behavioural tells, the code style, the language, the working rhythms, that attribution has come to rely on, which threatens to erode the one layer of evidence that false flags could not easily reach. And as more states build offensive capabilities, the background noise of plausible suspects grows, widening every probability band. The reasonable fear is not that attribution stops working, but that confident attribution becomes a luxury available only to the handful of governments with the intercepts and implants to see past the misdirection, while everyone else is left arguing over forgeries. Which returns to the uncomfortable core of it. Every time a country stands up and names its attacker, it is asking the rest of us to trust a probability we cannot fully audit. Sometimes that trust is earned by a leaked IP address or a building in Shanghai. Sometimes it rests on sources we will never see. The task is not to demand a certainty that the evidence can never provide, but to get honest about the odds, and to keep asking who is telling us, and why now.
Sources
- Thomas Rid and Ben Buchanan, Attributing Cyber Attacks, Journal of Strategic Studies, 2015.
- Office of the Director of National Intelligence, Intelligence Community Directive 203: Analytic Standards, 2007 (amended 2022).
- Sherman Kent, Words of Estimative Probability, CIA Studies in Intelligence, 1964.
- Kaspersky, Olympic Destroyer: who hacked the Olympics?, 2018.
- Securelist (Kaspersky), The Devil's in the Rich Header, 2018.
- US Department of Justice, Six Russian GRU Officers Charged in Connection with Worldwide Deployment of Destructive Malware, 2020.
- CyberScoop, U.S. and U.K. blame Russia for infamous 'NotPetya' cyberattacks, 2018.
- Mandiant, APT1: Exposing One of China's Cyber Espionage Units, 2013.
- US Department of Justice, U.S. Charges Five Chinese Military Hackers for Cyber Espionage, 2014.
- The Daily Beast, 'Lone DNC Hacker' Guccifer 2.0 Slipped Up and Revealed He Was a Russian Intelligence Officer, 2018.
- US Department of Justice, Indictment of 12 Russian GRU Officers (United States v. Netyksho), 2018.
- Time, FBI Accuses North Korea in Attack That Nixed 'The Interview', 2014.