Risk Limiting Audits: What They Are and Aren't
Philip Stark (Professor of Statistics · University of California at Berkeley)
Voting Village @ DEF CON 33 · Day 1 · Voting Village
Overview
Professor Philip Stark, a distinguished Professor of Statistics at the University of California at Berkeley and a former member of the Board of Advisors of the EAC, delivered a comprehensive and deeply technical presentation on Risk Limiting Audits (RLAs). This talk demystified the mathematical foundations and practical implications of RLAs, which have become recognized as the gold standard for post-election audits. Stark, widely credited as the intellectual and mathematical architect of RLAs, meticulously explained what these audits are designed to achieve, how they function at a fundamental level, and critically, what they are not capable of proving.

Key moments
- 0:00 Introduction of Philip Stark and origins of his work
- 3:00 Core principle: Present convincing evidence that reported winners won
- 4:00 Desirable security properties: software independence, contestability, defensibility
- 6:10 Defining Risk Limiting Audits (RLAs) and their purpose
- 7:20 What RLAs are NOT: Cannot make a poorly run election trustworthy
- 7:45 Historical context: First RLA papers and pilot implementation
Risk Limiting Audits: What They Are and Aren't
Speakers: Philip Stark, Professor of Statistics, University of California at Berkeley
Conference: Voting Village
YouTube: https://www.youtube.com/watch?v=f-QuFskAyOM
Overview
Professor Philip Stark, a distinguished Professor of Statistics at the University of California at Berkeley and a former member of the Board of Advisors of the EAC, delivered a comprehensive and deeply technical presentation on Risk Limiting Audits (RLAs). This talk demystified the mathematical foundations and practical implications of RLAs, which have become recognized as the gold standard for post-election audits. Stark, widely credited as the intellectual and mathematical architect of RLAs, meticulously explained what these audits are designed to achieve, how they function at a fundamental level, and critically, what they are not capable of proving.
The core objective of an RLA is to provide affirmative evidence that reported election winners genuinely won, or to trigger a full hand count if sufficient evidence is lacking. Stark emphasized that RLAs are a crucial ingredient in achieving evidence-based elections, ensuring that election outcomes are not merely announced but are demonstrably correct. His work, initiated almost two decades ago, has transformed the landscape of election integrity, offering a statistically rigorous method to verify results from potentially vulnerable voting technologies.
This article delves into the intricate details of Stark's presentation, outlining the historical context, the ingenious mathematical framework (including a unique "betting game" analogy), the practical advancements that make RLAs efficient, and the critical defensive implications for securing elections. It highlights the prerequisites for effective RLAs, stressing the paramount importance of a trustworthy paper trail and robust election administration practices.
Background
▶ Watch: Introduction of Philip Stark and origins of his work (0:00)
The journey to Risk Limiting Audits began serendipitously for Professor Philip Stark approximately 18 years ago, when he was invited by then-Secretary of State Debra Bowen to join the Post-Election Audit Standards Working Group (Peacewig). At the time, conventional post-election audits were largely inadequate. Statutory audits typically involved examining a fixed percentage of precincts or batches, retabulating votes within them, and comparing against machine tallies. States like Colorado, for instance, might simply re-tabulate 500 ballots from an unspecified source. Many states had no audits at all. These methods aimed to answer the question, "Are the machines working properly?" – a question Stark deemed problematic, as "properly" is not a binary state and some error is always expected. The more pertinent question, he argued, is "Did they work well enough in this election to find out who really won?"
Early computer scientists involved in election integrity, such as David Wagner (Stark’s first RLA co-author), also recognized the need for more principled audit designs. The prevailing "detection paradigm" of circa 2007 aimed for audits to have a large chance of finding at least one mistake if the outcome was wrong. However, Stark pointed out that in any large election, especially with hand-marked paper ballots, finding at least one mistake is almost guaranteed due to voter errors or machine misinterpretations. This paradigm failed to address the critical question: how big was the problem, and did it alter the outcome?
This realization led to the affirmative evidence paradigm, which underpins RLAs: "Has the audit given strong evidence that the reported winners really won? If not, collect more evidence or do a full hand count." This shift demanded a trustworthy paper trail, defined as a record that, if counted completely and accurately, would reveal the true winner. Stark cautioned that some vote records are "born untrustworthy" – malleable electronic records, or those produced by vulnerable technologies like Ballot Marking Devices (BMDs) or internet voting systems, which can compromise the paper trail itself. Even with hand-marked paper ballots, maintaining trustworthiness requires physical inventories, demonstrably secure chain of custody, physical security, audits of eligibility, and ballot accounting.
Crucially, RLAs require a known upper bound on the number of cards cast in each contest, typically derived from voter participation records rather than merely registered voters. The importance of RLAs was formally recognized in 2018 when the National Academies of Sciences, Engineering, and Medicine's report on U.S. elections issued two primary recommendations: use paper ballots and conduct Risk Limiting Audits. Today, laws authorizing RLAs exist in about 15 states, although their implementation quality varies significantly.
Key Findings
▶ Watch: Desirable security properties: software independence, contestability, defensi... (4:00)
Professor Stark's presentation distilled the essence of Risk Limiting Audits into several core findings and principles:
- Definition and Purpose: An RLA is fundamentally "any procedure that limits the risk that you will certify an incorrect electoral outcome." The "risk" is defined as the largest chance that a wrong outcome will be certified. RLAs are designed to correct wrong outcomes with high probability and, importantly, "will never make a right outcome wrong." Their primary goal is to provide affirmative evidence that reported winners truly won, or to compel a full hand count if that evidence is insufficient, before results are certified.
- "Garbage In, Garbage Out": A critical caveat is that RLAs cannot retroactively make a poorly run election trustworthy. Stark unequivocally stated, "It's garbage in, garbage out." Applying RLA procedures to untrustworthy paper trails or poorly managed elections is a form of "disinformation" that distracts from the foundational work required to ensure a secure and verifiable election.
- Practicality and Efficiency: Contrary to common misconceptions that RLAs are too resource-intensive for large elections or numerous contests, Stark demonstrated their high efficiency. Modern RLA methods, particularly those employing targeted sampling and consistent sampling, allow for auditing every contest in an election, even in large jurisdictions with hundreds of contests. For example, Orange County, California, audited 181 contests in 2020 by examining only 0.65% of cast cards, and 214 contests in 2022 with 3.1% of cards (including some very close margins). This makes full hand counts for most contests unnecessary.
- Not Tabulation Audits: A significant clarification offered by Stark is that RLAs are not tabulation audits. They do not check whether the machine tabulation is correct; they only check "who won." He provided an example where every single ballot was mistabulated by a machine, yet the correct winner was declared. An RLA would correctly stop without a full hand count in such a scenario, confirming the winner despite tabulation errors. This directly refutes claims by some election officials that RLAs prove "no votes were flipped."
- Protection Against Worst-Case Errors: RLAs provide robust protection against errors regardless of their distribution. It doesn't matter if errors are clustered in one precinct or randomly dispersed; the RLA methodology accounts for the worst-case scenario to ensure a high probability of detecting and correcting a wrong outcome.
- Adaptability to Social Choice Functions: RLA methods have been developed for a wide array of social choice functions, including plurality, multi-winner plurality, supermajority, proportional representation schemes (like D'Hondt), STAR voting, and Instant Runoff Voting (IRV). While some complex systems like Single Transferable Vote (STV) still present challenges, significant progress has been made.
- Key Trust Assumptions: The efficacy of RLAs hinges on a few, but crucial, trust assumptions:
- A known upper bound on the number of cards cast in each contest.
- A commitment to a set of immutable card identifiers before the audit begins. These identifiers, ideally printed indelibly on each ballot, ensure transparency and verifiability during ballot retrieval. Stark noted that auditors need to be able to request "card number 2719" and observers must know the correct ballot has been retrieved. However, the system can "fail up" – if identifiers are misused, it only increases the measured risk, potentially leading to an unnecessary hand count but not undermining the risk limit.
- The audit must involve physical inspection of paper ballots by human eyeballs, not just digital images, due to the malleability of digital data and the risk of undetected scanning errors (e.g., missing or duplicated batches).
Technical Deep Dive
▶ Watch: Defining Risk Limiting Audits (RLAs) and their purpose (6:10)
The core of Professor Stark's explanation of Risk Limiting Audits revolves around an intuitive yet statistically rigorous betting game analogy. This approach, detailed in the "alpha" paper (Stark et al., 2012), unifies RLA methods across various election types.
The Betting Game Analogy:
Imagine a simple two-candidate plurality contest between Alice and Bob, where Alice is the reported winner. You start with a bankroll of one dollar and repeatedly pull a ballot at random.
- If the ballot shows a vote for Alice, you win your bet back (doubling your money).
- If it shows a vote for Bob, you lose your bet.
- If it's for another candidate (Carol) or a non-vote, no money changes hands.
The rules: you can bet up to your current bankroll, you can't go into debt, and if you go broke, you must perform a full hand count.
The statistical power of this game comes from Jean V's 1939 paper, which states that if Alice did not actually win (i.e., did not get more votes than Bob), the chance that you will ever multiply your initial bankroll by a factor of K is at most 1/K. For instance, if you manage to turn your initial dollar into $100, there's at most a 1% chance this would happen if Alice didn't truly win. This provides strong evidence that Alice really did beat Bob. If you achieve a bankroll multiplication of 20 (e.g., $1 to $20), you've effectively conducted an RLA at a 5% risk limit.
Generalizing the Betting Game:
Stark demonstrated how this principle extends to more complex scenarios:
- Multi-Winner Plurality: For contests where multiple candidates win (e.g., "vote for up to five"), a separate betting game is set up for every reported winner against every reported loser. These games are played in parallel. If you win a sufficient amount (e.g., $20) in all parallel games, the audit stops, affirming all reported winners.
- Encoding Votes as Numbers: To generalize further, votes are encoded as numbers. In an Alice vs. Bob contest:
- A vote for Alice becomes
1. - A vote for Bob becomes
0. - A vote for any other candidate or a non-vote becomes
0.5.
If Alice truly won, the average of this list of numbers across all ballots will be greater than 0.5. The betting game then becomes about whether the average of this list is greater than 0.5.
- Abstract Payoff Formula: The payoff for a bet
Bon a ballotX(whereXis the numerical encoding of the vote) is2 B (X - 0.5). Stark meticulously walked through how this formula perfectly encapsulates the Alice/Bob/Carol payoff rules:
- If
X = 1(Alice), payoff is2 B (1 - 0.5) = 2 B 0.5 = B. You win your bet back. - If
X = 0(Bob), payoff is2 B (0 - 0.5) = 2 B (-0.5) = -B. You lose your bet. - If
X = 0.5(Carol/non-vote), payoff is2 B (0.5 - 0.5) = 0. No money changes hands.
This abstract formulation allows the same statistical principles to apply to any social choice function, where different encoding schemes for X are used.
Optimal Betting and Efficiency:
The speed and efficiency of an RLA depend on how much an auditor bets. Betting the entire bankroll each time would be fastest (e.g., five Alice votes to reach $32), but also carries a high risk of going broke if a Bob vote is drawn. The goal is to grow the bankroll as rapidly as possible while minimizing the risk of going broke.
- Kelly Criterion (1956): John Kelly's work in information theory established the optimal betting strategy when the odds are known, maximizing the long-term growth rate of capital.
- Adaptive Strategies: Since the true odds (the actual vote proportions) are unknown at the start of an audit, modern RLA methods use adaptive strategies. These methods, influenced by Abraham Wald's 1945 work on sequential testing of statistical hypotheses, start by betting as if the reported results are correct, but then continuously learn from the sampled ballots and adjust the betting strategy over time. This "shrinkage" approach allows for near-optimal performance without prior knowledge of the true outcome.
Comparison Audits and Variance Reduction:
Stark also explained how comparison audits, which leverage cast vote records (CVRs) or other reported values, achieve greater efficiency. Instead of the raw vote X, the numbers in the list for comparison audits are derived from 1 - (reported_value - original_value) / (2 * reported_margin).
If the CVRs are perfectly accurate, the reported_value will equal the original_value, making the numerator zero. This results in a list where every element is a constant, 1 - 0 / (2 * reported_margin) = 1. Betting on a list of all 1s is a "sure thing bet." The auditor can bet a much larger portion of their bankroll without risk, leading to a much faster audit. The magic here is the reduction in variance: the less uncertainty in the underlying data (i.e., the closer the CVR matches the actual ballot), the more aggressively you can bet, and the faster you reach the desired risk limit. This also enables hybrid audits, where batch-level subtotals can be used as reference values for individual cards within that batch, maintaining statistical validity.
Demo / Proof of Concept
▶ Watch: What RLAs are NOT: Cannot make a poorly run election trustworthy (7:20)
While Professor Stark's main presentation was a deep dive into the theoretical and mathematical underpinnings of Risk Limiting Audits, he concluded by announcing a separate, participatory demonstration to be held immediately following the talk in another room. This hands-on session was described as involving "fake ballots and some dice," where attendees would have the opportunity to actively audit a mock election using the principles he had just explained. This practical demonstration was intended to solidify the conceptual understanding of the betting game analogy and the mechanics of an RLA in a tangible, interactive format, allowing participants to experience the process firsthand rather than just observing a pre-recorded demo.
Defensive Implications
▶ Watch: Historical context: First RLA papers and pilot implementation (7:45)
Professor Stark's talk provided critical insights for election defenders, highlighting both the power of Risk Limiting Audits and the essential preconditions for their effectiveness:
- Prioritize a Trustworthy Paper Trail: The foundational defensive measure is establishing and maintaining a trustworthy paper trail. This means advocating for hand-marked paper ballots, as they are less susceptible to systemic hacking than electronic systems. Crucially, the paper trail must be accompanied by robust administrative practices: physical inventories of all ballots, a demonstrably secure chain of custody from voter to audit, physical security of all election materials, rigorous eligibility audits, and comprehensive ballot accounting. Stark emphatically stated this is "Problem One" – without it, RLAs are compromised.
- Beware of "Magic Dust" and Disinformation: Defenders must be vigilant against the misuse of RLAs as a "fig leaf" for poorly run elections or untrustworthy systems. Applying RLA procedures to malleable digital images (without corresponding physical paper), or to systems that lack proper chain of custody, is ineffective and can be a form of disinformation. RLAs do not validate the accuracy of the underlying tabulation or prove that "no votes were flipped"; they merely verify the correctness of the declared winners. Claims to the contrary should be challenged.
- Insist on Immutable Ballot Identifiers: For an RLA to be publicly verifiable, each ballot needs an indelible, immutable identifier. Relying on sequential numbering within batches (e.g., "the 17th ballot in batch X") is insufficient as it's not truly immutable and can be manipulated. Defenders should push for direct printing of unique, public identifiers on ballots. Stark referenced the "Nonsuch paper" which details how this can be done while preserving voter privacy. This ensures that when a specific ballot is selected for audit, observers can confirm the correct physical ballot has been retrieved.
- Demand Physical Ballot Sampling: The audit must involve eyeballs on paper. Relying solely on digital ballot images for comparison is a significant vulnerability. Scanners are computers and can be compromised, leading to issues like missing batches, duplicated batches, or altered images – errors that would likely not be detected by an image-based audit alone. Physical sampling provides the necessary independent verification.
- Understand RLA Scope and Limitations: Defenders should educate themselves and the public that RLAs confirm the reported winners, not the absence of all errors or the perfect accuracy of tabulation. This nuanced understanding prevents overstating the audit's capabilities and focuses efforts on areas where real improvements are needed.
- Advocate for Proper Randomization: The integrity of the sampling process is paramount. Randomization for ballot selection should be robust and publicly verifiable, typically involving a public ritual (like rolling 10-sided dice) to generate a seed for a high-quality cryptographically secure pseudo-random number generator (e.g., based on SHA-256 in counter mode). This minimizes the risk of manipulated samples.
- Push for Effective RLA Implementation: While many states have RLA laws, their actual implementation varies widely. Defenders should advocate for robust, statistically sound RLA practices, including targeted sampling and consistent sampling techniques, which make auditing every contest practical and efficient. This also means understanding that different contests may require different sampling rates based on their margin.
In essence, Professor Stark's message to defenders is to focus on the fundamentals of election integrity – secure paper ballots, transparent processes, and verifiable records – and then to leverage RLAs as the powerful, statistically sound tool they are, rather than a superficial fix for deeper systemic issues.
Key Takeaways
- RLAs are the Gold Standard for Verifying Winners: Risk Limiting Audits provide strong statistical evidence that the reported winners of an election truly won, or they mandate a full hand count if evidence is insufficient to meet a predefined risk limit (e.g., 5%).
- The Betting Game Analogy Simplifies Complex Statistics: The core mathematical principle of RLAs can be understood through a "betting game" where auditors "win money" if the reported outcome is correct. The probability of winning significant amounts is extremely low if the reported outcome is wrong, providing compelling evidence.
- RLAs Are Not Tabulation Audits: A critical distinction is that RLAs verify who won, not whether the machine tabulation was perfectly accurate. They do not prove that "no votes were flipped" or that all machines worked correctly, only that the final outcome aligns with the ballots.
- Trustworthy Paper Trails Are Non-Negotiable: Effective RLAs absolutely depend on a physically secure, transparent, and trustworthy paper trail, complete with immutable ballot identifiers, robust chain of custody, and ballot accounting. RLAs cannot fix a fundamentally flawed or untrustworthy election system.
- Efficiency Makes Auditing All Contests Practical: Advanced RLA methods, particularly targeted sampling and consistent sampling, make it highly efficient and practical to audit every contest in large elections, often requiring inspection of only a small percentage of ballots overall.
- Physical Inspection is Paramount: For true security, auditors must physically inspect paper ballots. Relying solely on digital images is insufficient due to the malleability of digital data and the risk of undetected errors like missing or duplicated batches.
About the Speaker(s)
Professor Philip Stark is a distinguished Professor of Statistics at the University of California at Berkeley. His journey into election integrity began almost by accident 18 years ago, when a cold call from David Wagner, then-chairman of the Post-Election Audit Standards Working Group (Peacewig), invited him to join. That call, as the introduction quipped, "ruined his life" by dedicating a huge fraction of his time and research energy to this critical field.
Stark is widely recognized as the intellectual and mathematical founder of Risk Limiting Audits (RLAs) and the broader concept of evidence-based elections. Beyond developing the mathematical theory, he has tirelessly traveled the country, simplifying and refining RLA techniques, educating election officials, running RLA trials, and serving as an expert witness, notably in the Curling v. Raffensperger trial in Georgia. He has collaborated with numerous co-authors, including computer scientists like David Wagner and Andrew Appel, and election officials, ensuring his work is both theoretically sound and practically implementable. His pioneering efforts have almost single-handedly created and dominated the field of RLAs, which are now considered the gold standard for election auditing.
Reviews
Dr. Zero (Offensive Security Researcher) — STRONG ACCEPT
Stark is the originator of this field — not a popularizer, not a secondary contributor — and it shows. The mathematical core is explained with genuine depth, the distinctions between what RLAs do and don't prove are precisely drawn, and the practical limitations get as much airtime as the methodology's strengths. A few sections read like transcript padding rather than technical exposition, and the 'demo' is essentially deferred to another room, but neither knocks this below a strong accept.
Heather Calloway (CISO) — SOLID
Stark is the right person to explain RLAs and he does it well — the betting game analogy is genuinely clarifying, and the distinction between tabulation audits and outcome verification is one that many election officials get wrong. But this is a deep technical education session, not a governance or operational decision talk, and the audience it will change the most is statisticians and technically curious attendees, not the administrators and policymakers who control whether any of this gets implemented correctly.