Game-theoretic Statistics and Sequential Anytime-Valid Inference: Game-theoretic Statistics and Sequential Anytime-Valid Inference (SAVI): A Martingale Theory of Evidence
Aaditya Ramdas
International Conference on Machine Learning 2025 · Tutorial
Overview
Aaditya Ramdas’s tutorial at ICML 2025 introduced attendees to the rapidly evolving field of game-theoretic statistics and Sequential Anytime-Valid Inference (SAVI), presenting it as a foundational shift in how we approach statistical inference. Ramdas, a distinguished statistician and recipient of numerous prestigious awards, posits that this framework, rooted in a martingale theory of evidence, offers a robust, flexible, and universal alternative to classical statistical methods, particularly addressing the pervasive issues associated with p-values. The talk highlighted how concepts from gambling and betting provide a powerful lens through which to unify and re-envision hypothesis testing and estimation, enabling valid and efficient inference even in complex, adaptive, and sequential data analysis scenarios.

Key moments
- 0:00 Introduction to Game-theoretic Statistics and E-values
- 2:00 Gambling as the foundation of game-theoretic statistics
- 2:40 Tutorial outline: Hypothesis testing and estimation
- 3:30 Motivation for Sequential Anytime-Valid Inference (SAVI)
- 4:30 Karl Popper: Why statisticians test hypotheses
- 5:20 Formally defining e-values and e-processes
Game-theoretic Statistics and Sequential Anytime-Valid Inference (SAVI): A Martingale Theory of Evidence
Speakers: Aaditya Ramdas
Conference: ICML 2025
YouTube: https://slideslive.com/39043357
Overview
Aaditya Ramdas’s tutorial at ICML 2025 introduced attendees to the rapidly evolving field of game-theoretic statistics and Sequential Anytime-Valid Inference (SAVI), presenting it as a foundational shift in how we approach statistical inference. Ramdas, a distinguished statistician and recipient of numerous prestigious awards, posits that this framework, rooted in a martingale theory of evidence, offers a robust, flexible, and universal alternative to classical statistical methods, particularly addressing the pervasive issues associated with p-values. The talk highlighted how concepts from gambling and betting provide a powerful lens through which to unify and re-envision hypothesis testing and estimation, enabling valid and efficient inference even in complex, adaptive, and sequential data analysis scenarios.
The core premise of game-theoretic statistics is to transform statistical problems into betting games. By designing games where a rational "player" (statistician) bets against a null hypothesis, the player’s accumulated wealth directly quantifies the evidence against that null. This approach inherently accommodates continuous data monitoring and adaptive decision-making—practices that notoriously invalidate classical p-value-based inference—without compromising the validity of statistical claims. Ramdas demonstrated how this paradigm not only provides a principled solution to long-standing challenges like p-hacking and "sampling to a foregone conclusion" but also yields novel, asymptotically optimal strategies for a wide range of problems, from mean estimation to complex causal inference in observational studies.
This tutorial is significant because it challenges the conventional statistical education, urging a re-learning of inference from first principles. By grounding inference in game theory, it offers a universal language for testing any empirically testable hypothesis and deriving estimators with strong, time-uniform guarantees. The practical implications are profound, extending to fields like AB testing, clinical trials, and off-policy evaluation, where real-world applications at major tech companies are already leveraging SAVI methods. Ramdas’s vision positions game-theoretic statistics not just as an academic curiosity, but as a pragmatic and indispensable tool for the future of data-driven decision-making.
Background
▶ Watch: Introduction to Game-theoretic Statistics and E-values (0:00)
The motivation for game-theoretic statistics stems from long-standing problems and misuses of classical statistical inference, particularly the p-value. As eloquently articulated by Karl Popper, scientific theories must be empirically testable. In statistics, this translates to formulating a null hypothesis (e.g., "no effect," "nothing interesting is going on") and attempting to contradict it using data—a process often described as "stochastic proof by contradiction." However, the conventional p-value, which quantifies the probability of observing data as extreme as, or more extreme than, what was observed, given the null hypothesis is true, is fraught with issues.
Ramdas highlighted the infamous "power posing controversy" as a prime example of p-value misuse. Researchers, in an attempt to find statistically significant results, engaged in "sampling to a foregone conclusion" or p-hacking: continuously collecting and analyzing data, stopping only when a p-value fell below the arbitrary 0.05 threshold. This practice, while common, drastically inflates the false positive rate (Type I error), leading to spurious discoveries. Even after decades of warnings, practitioners continue to "peek" at their data, demonstrating a fundamental tension between the desire for adaptive experimentation and the fixed-sample assumptions of classical p-values. Similar issues plague confidence intervals, where continuous monitoring can lead to "loopy and false behavior," with intervals losing their stated coverage guarantees.
The solution proposed by game-theoretic statistics is Sequential Anytime-Valid Inference (SAVI). SAVI methods are designed to "legalize" continuous monitoring and allow for adaptive decisions—halting or continuing an experiment for any reason, at any time, without violating the validity of the claims. This flexibility is crucial in exploratory studies, industry AB testing, and neuroscience, where data collection is expensive and the "right" sample size is often unknown in advance. The historical roots of this idea trace back to early 20th-century work by Ville (1939) and later formalized by Vovk and Shafer in their seminal book on game-theoretic probability. Ramdas’s work builds upon this foundation, adapting the principles of gambling and martingales to create a universal and robust framework for statistical inference.
Key Findings
▶ Watch: Tutorial outline: Hypothesis testing and estimation (2:40)
The central contributions of game-theoretic statistics, as presented by Ramdas, revolve around redefining evidence and inference through the lens of betting games. Key findings include:
- E-values as a direct measure of evidence: Unlike p-values, which are probabilities under the null, e-values directly quantify evidence against the null. A larger e-value signifies stronger evidence. If the null hypothesis is true, the expected value of an e-value is less than or equal to one.
- E-processes and Test Martingales as Fundamental Objects: The sequential analogue of e-values are e-processes, which are sequences of e-values valid at any stopping time. A specific type of e-process, the test supermartingale (or test martingale), where the expected wealth under the null decreases (or stays constant) over time, forms the backbone for sequential hypothesis testing. This property holds simultaneously for every distribution within the null hypothesis, a crucial distinction from classical martingales.
- Ville's Inequality for Anytime-Valid Inference: This fundamental inequality guarantees that if an e-process's wealth ever crosses a threshold of
1/alpha, the probability of this occurring under the null is at mostalpha. This provides a time-uniform Markov's inequality for non-negative supermartingales, enabling sequential testing with controlled false positive rates at any arbitrary stopping time. - Universality of E-processes for Sequential Testing: Ramdas presented a remarkable theorem: every level
alphasequential test for any null hypothesis can be recovered by thresholding an e-process at1/alpha. This establishes e-processes as necessary and sufficient for sequential testing, making the game-theoretic framework universal. - Log-Optimality and Kelly Betting: For favorable betting games, the Kelly criterion dictates betting a fraction of wealth that maximizes the expected logarithm of wealth. This strategy, yielding exponential wealth growth, is proven to be asymptotically optimal and establishes a clear connection between optimal gambling and Kullback-Leibler (KL) divergence (relative entropy), as shown by Kelly and later generalized by Breiman.
- Strong Duality for Optimal Betting Strategies: For composite nulls against simple alternatives, the optimal one-round bet (the numeraire, X*) is a specific likelihood ratio of the alternative (Q) to a special measure (P*) called the reverse information projection. This numeraire is log-optimal, and its expected log wealth growth rate under the alternative is precisely the KL divergence between Q and P*, establishing a powerful strong duality.
- Confidence Sequences (CS) for Estimation: As a dual to sequential testing, CS provide a sequence of confidence intervals that offer simultaneous coverage guarantees for a parameter
thetaacross all time pointst. This means the probability thatthetalies within the interval[Lt, Ut]for alltfrom 1 to infinity is at least1-alpha, resolving the issue of miscoverage in continuous monitoring. - Time-Uniform Central Limit Theory (TU-CLT) and Asymptotic Confidence Sequences (AsymCS): For settings where non-asymptotic guarantees are impossible (e.g., unknown finite variance, non-IID data), Ramdas introduced AsymCS. An AsymCS is an arbitrarily precise, almost sure approximation to a non-asymptotic CS for large
t. The TU-CLT provides a sequential analogue to the classical CLT, allowing the construction of AsymCS for mean estimation and, critically, for doubly robust (DR) causal inference in observational settings.
Technical Deep Dive
▶ Watch: Motivation for Sequential Anytime-Valid Inference (SAVI) (3:30)
Game-theoretic statistics fundamentally reframes statistical inference as a betting game. The core idea is to design a game where a player (the statistician) bets against a null hypothesis (H0). If H0 is true, the player should not be able to systematically make money. If H0 is false, an astute player should be able to gain wealth.
The foundational element is the wealth process, denoted L_t. Starting with an initial capital L_0 = 1, the wealth at time t is a product: L_t = product(i=1 to t) (1 + lambda_i * R_i), where lambda_i is the fraction of current wealth bet, and R_i is the outcome (e.g., +1 for win, -1 for loss). The bets lambda_i must be predictable, meaning they are based on all past observations and the current state, but not on the outcome of the current bet itself.
The crucial property under the null hypothesis is that the wealth process L_t must be a non-negative martingale or supermartingale. A martingale implies E[L_t | F_{t-1}] = L_{t-1}, meaning the expected wealth, given all past information (F_{t-1}), remains constant. A supermartingale implies E[L_t | F_{t-1}] <= L_{t-1}, meaning expected wealth can only decrease or stay constant. This property must hold simultaneously for all distributions P belonging to the null hypothesis set script P. This "simultaneity" is what distinguishes game-theoretic martingales from classical ones.
From this, the concept of e-values emerges. An e-value E for testing a null script P is a non-negative random variable such that E_P[E] <= 1 for every P in script P. An e-process is a sequence of e-values E_t such that E_P[E_tau] <= 1 for any stopping time tau. This is critical for anytime-valid inference, as it means the expected wealth cannot exceed the initial capital even if the experiment stops at a data-dependent time.
Ville's inequality is a cornerstone, stating that the probability of an e-process L_t ever exceeding 1/alpha (i.e., P(sup_t L_t > 1/alpha)) is at most alpha, under the null. This directly yields a level alpha sequential test: reject the null the first time L_t crosses 1/alpha.
The design of optimal betting strategies is key to the efficiency of this framework. For simple nulls and simple alternatives (P vs. Q), the optimal bet in each round is the likelihood ratio Q(X_t) / P(X_t). This process, L_t = product(i=1 to t) (Q(X_i) / P(X_i)), is a classical likelihood ratio process, which is a test martingale under P. The expected log wealth growth rate under the alternative Q is the KL divergence D_KL(Q || P). This is the essence of Kelly betting and log-optimality, where maximizing expected log-wealth maximizes the exponential growth rate of capital.
For composite nulls (arbitrary set script P) against a simple alternative (Q), the optimal one-round bet, termed the **numeraire (X*)**, is a likelihood ratio of Q to a special measure P*. This P* is the reverse information projection of Q onto the bipolar of script P, representing the "closest" distribution to Q within the effective null. This strategy is log-optimal, and a strong duality theorem connects the numeraire's expected log-growth rate to inf_P in bipolar(P) D_KL(Q || P). When dealing with composite alternatives, a mixture strategy (hedging) over these alternatives is employed, effectively averaging individual betting strategies. This mixture strategy adaptively achieves optimal growth rates, even without a priori knowledge of the true alternative.
For estimation, the framework yields confidence sequences (CS). A CS is a sequence of intervals [L_t, U_t] such that P(forall t, theta in [L_t, U_t]) >= 1-alpha. These are derived by inverting families of sequential tests. For instance, to estimate a mean mu, one plays a continuum of games, each testing a different null mu = m. The confidence set C_t is then the set of all m for which the wealth in the m-th game has not yet crossed 1/alpha. The GRAPPA (Growth Rate Adaptive to Particular Alternative) strategy for mean estimation uses a generalized Kelly bet, lambda_mt = (sample_mean - m) / sample_variance, adapting to unknown variance and providing tighter bounds than classical methods.
Finally, for problems requiring asymptotic theory, Asymptotic Confidence Sequences (AsymCS) were introduced. An AsymCS [mu_hat_t +/- b_t] is an almost-sure approximation to a non-asymptotic CS [mu_hat_t +/- b*_t] for large t. The Time-Uniform Central Limit Theory (TU-CLT) provides the theoretical basis for constructing AsymCS under minimal assumptions (e.g., finite unknown variance), extending the power of the CLT to sequential settings. This is particularly relevant for doubly robust (DR) causal inference in observational studies, where an AsymCS for the average treatment effect can be constructed using sequentially cross-fitted DR estimators, provided the product of errors in estimating nuisance parameters (regression function and propensity score) is sufficiently small (small o(sqrt(log t / t))).
Experimental Setup & Results
▶ Watch: Karl Popper: Why statisticians test hypotheses (4:30)
Ramdas illustrated the power of game-theoretic statistics through several compelling examples and empirical comparisons.
A foundational example was the "Lady Tasting Tea" experiment, historically analyzed by R.A. Fisher. Muriel Bristol correctly identified 8 out of 8 cups, yielding a p-value of 1/70. Ramdas re-enacted this with his wife, Leila, in the "Lady Keeps Tasting Coffee" experiment, betting on whether espresso was added to milk or vice-versa. Leila started with one lira and placed predictable bets. If the null (no difference in preparations) were true, her expected wealth would remain constant (a fair game). However, if she could discern a difference, her wealth would grow. This simple setup demonstrated how accumulated wealth directly measures evidence against the null.
For mean estimation of bounded (0,1) IID random variables, Ramdas compared confidence sequences (CS) derived from betting strategies against classical Hoeffding and empirical Bernstein-style confidence intervals.
- Classical intervals: When estimating the mean of a Gaussian distribution, classical confidence intervals (e.g.,
+/- z_{1-alpha} / sqrt(n)) were shown to miscover the true mean infinitely often with probability one in continuous monitoring scenarios. - Game-theoretic CS: The proposed CS (e.g.,
mu_hat_t +/- sqrt(log(1/alpha) + log log t / t)) were wider than classical intervals but guaranteed simultaneous coverage for alltwith high probability (1-alpha). Visualizations showed CS bands consistently covering the true mean, while classical intervals frequently failed. - Efficiency: The GRAPPA betting strategy, an approximate Kelly bet (
lambda_mt = (sample_mean - m) / sample_variance), and a mixture strategy over possiblelambdavalues, yielded significantly tighter CS compared to existing empirical Bernstein intervals (e.g., Maurer and Pontil 2009, Audibert et al 2007). The article also mentioned a novel closed-form CS that is asymptotically sharp, matching Bernstein's width without requiring prior knowledge of the variance. - Generalization: This approach extended to quantile estimation (e.g., for median,
empirical_median +/- sqrt(log(1/alpha) + log log t / t)) and even simultaneous coverage of all quantiles over time (time-uniform DKW inequality). Similarly, it was shown to apply to sequential covariance matrix estimation, forming ellipsoids that cover the true matrix over time.
For kernel two-sample testing, the game-theoretic framework was used to derive a sequential version of the Maximum Mean Discrepancy (MMD). This process adapts to unknown differences between two distributions P and Q, with its wealth growth rate proportional to the MMD, eliminating the need to pre-specify a sample size n.
In doubly robust (DR) causal inference in observational settings, Ramdas presented asymptotic confidence sequences (AsymCS) for the average treatment effect.
- Visual comparisons: Plots illustrated how unadjusted or parametrically adjusted estimators exhibited asymptotic bias, failing to cover the true average treatment effect. In contrast, non-parametric adjustments (e.g., using super learners or stacking for regression functions and propensity scores) successfully removed bias asymptotically, yielding valid and tighter AsymCS.
- Time-varying effects: The framework was shown to extend to covering time-varying treatment effects, demonstrating its robustness beyond IID assumptions.
- Real-world Deployment: The practical utility of these methods was underscored by their deployment in production systems at major companies: Amazon and Netflix use game-theoretic ideas for AB testing, while Adobe and Microsoft employ confidence sequences for off-policy evaluation in bandits. GrowthBook, a Y Combinator startup, also utilizes AsymCS.
Practical Implications
▶ Watch: Formally defining e-values and e-processes (5:20)
The game-theoretic statistics framework, with its emphasis on Sequential Anytime-Valid Inference (SAVI), carries profound practical implications for practitioners across various fields:
- Adaptive Experimentation and Continuous Monitoring: For practitioners in industry (e.g., AB testing, online learning) or research (e.g., clinical trials, neuroscience), SAVI legalizes continuous monitoring and adaptive decision-making. Researchers can peek at their data, stop experiments early, or extend them without invalidating their statistical claims, a critical advantage over classical p-value or confidence interval methods that break down under such practices. This provides immense flexibility and efficiency, allowing teams to react to data as it arrives.
- Robustness to P-hacking and Misuse: By replacing p-values with e-values and e-processes, the framework naturally mitigates the risks of p-hacking and "sampling to a foregone conclusion." The evidence (wealth) accumulates directly, and its interpretation is transparent and consistent regardless of when an experiment is stopped.
- Enhanced Statistical Efficiency: Optimal betting strategies, particularly those leveraging log-optimality and connections to KL divergence, often yield tighter confidence sequences and more powerful tests compared to classical approaches. This translates to needing less data to achieve a desired level of certainty, or achieving higher certainty with the same amount of data. For instance, empirical comparisons showed game-theoretic confidence sequences to be much tighter than Hoeffding or empirical Bernstein intervals for mean estimation.
- Universality and Generalizability: The framework provides a universal language for hypothesis testing and estimation, applicable across diverse data structures (vectors, matrices, general distributions) and assumptions (IID, non-IID, martingale settings). This means a unified approach can be used for problems ranging from simple mean estimation to complex non-parametric two-sample testing or doubly robust causal inference.
- Principled Causal Inference in Observational Studies: The development of Asymptotic Confidence Sequences (AsymCS) and the Time-Uniform Central Limit Theory (TU-CLT) enables robust sequential causal inference in observational settings. This allows infra teams and model builders to make valid, adaptive inferences about treatment effects, even when estimating nuisance parameters (like propensity scores and regression functions) on the fly. However, it's crucial to acknowledge that this still requires assumptions on the absence of unobserved confounding and the accurate estimation of nuisance parameters.
- Quantifying Evidence Directly: E-values offer a more intuitive measure of evidence than p-values. An e-value of 20 means there's 20 times more evidence against the null than if the null were true, allowing for a continuous scale of evidence rather than a binary "reject/fail to reject" decision.
- Tradeoffs and Limitations: While powerful, the framework is not without its nuances. The choice of "game" or filtration (which data to use) for a completely new non-parametric problem remains a design element, not yet fully automatized. Furthermore, while AsymCS are robust, their finite-sample behavior (rate of convergence to asymptotic validity) still depends on underlying data properties and can be a subject of further quantification, similar to classical asymptotic methods. The computational cost of discretizing a continuum of games (e.g., for mean estimation) is also a practical consideration, though often deemed manageable with modern computing resources.
Key Takeaways
- Game-theoretic statistics provides a universal and robust framework for statistical inference, transforming hypothesis testing and estimation into betting games to address limitations of classical methods like p-values.
- E-values and E-processes are fundamental for sequential, anytime-valid hypothesis testing, directly quantifying evidence against the null hypothesis and enabling continuous monitoring without compromising validity.
- Confidence Sequences (CS) offer simultaneous coverage guarantees for estimation across all time points, providing a principled solution for adaptive experiments where traditional confidence intervals fail due to data peeking.
- Optimal betting strategies, such as Kelly betting (log-optimality), are rooted in likelihood ratios and KL divergence, leading to highly efficient inference methods that can adapt to unknown data distributions and yield tighter bounds than classical approaches.
- The framework extends to complex problems like sequential doubly robust causal inference in observational settings, leveraging Asymptotic Confidence Sequences (AsymCS) and Time-Uniform Central Limit Theory (TU-CLT) for valid adaptive experimentation.
- Practical applications are already being deployed in industry, demonstrating the real-world utility of these methods for AB testing, off-policy evaluation, and other data-driven decision-making scenarios.
About the Speaker(s)
Aaditya Ramdas is a distinguished figure in the field of statistics and machine learning. He holds a professorship and is widely recognized for his groundbreaking contributions to game-theoretic statistics. His work in this area is considered to be at the forefront of the future of statistics, particularly in advancing concepts like e-values. Ramdas has received numerous accolades for his research, including the Presidential Early Career Award, which is the highest distinction awarded by the U.S. government for young scientists. He is also a Kavli Fellow, a Sloan Fellow, and an Emerging Leader Award recipient from the Committee of Presidents of Statistical Societies (COPSS). Furthermore, he has been elected as a Fellow of the Institute of Mathematical Statistics (IMS) and was recognized as Statistician of the Year 2025. His research interests span game-theoretic statistics, connections to online learning, and information theory, all of which were intricately woven into this tutorial.
Reviews
Maya Iyer (Theoretical ML Researcher) — STRONG ACCEPT
Ramdas delivers a rigorous and well-organized tutorial on game-theoretic statistics and Sequential Anytime-Valid Inference, grounding the framework in martingale theory and making a credible case that e-values and e-processes are the right objects for sequential hypothesis testing. The theoretical backbone is sound — Ville's inequality, the universality theorem for e-processes, Kelly/log-optimality, the reverse information projection, and the time-uniform CLT are all presented with appropriate precision. The talk does not pretend to be a paper with original proofs; it is a tutorial, and judged as such it is unusually strong: it builds from first principles, situates the work honestly in…
Chen Zhao (Applied ML Researcher & Empiricist) — STRONG ACCEPT
Ramdas delivers a technically rigorous and intellectually coherent tutorial on game-theoretic statistics and sequential anytime-valid inference, grounding a decade of his own research program in first principles and connecting it to concrete industrial deployment. The framework is genuinely significant — e-values and confidence sequences solve a real problem that classical p-values cannot, and the universality theorem (every level-alpha sequential test is recovered by thresholding an e-process) is a strong result. The talk earns a 4 rather than a 5 primarily because it is a tutorial synthesizing an existing body of work rather than presenting a single new experimental contribution, and the…
→ Top-rated talks at International Conference on Machine Learning 2025
All talks from International Conference on Machine Learning 2025