Appendix: The Method
“An attributed quotation is evidence only of what its author said. It is not evidence that the rule works.”
This appendix describes the machinery behind the conclusions in the main book. It can be skipped without losing the argument — but if you want to know why some claims survive and others collapse, this is the key.
The evidence chain
Most personal finance proceeds by assertion: famous investor says X, therefore X. The problem is that famous investors disagree. We need something more discriminating than authority.
Every claim in this project was run through a six-gate chain:
claim → mechanism → evidence/counterevidence → boundary/failure mode → modern relevance → disposition
A claim that fails at any gate does not become a rule. Let’s walk through each gate.
Gate 1: The claim itself
First, we ask: what would an investor actually do differently?
“Bonds are safe” is not an actionable claim. A rule must be specific enough to falsify: “An investor should hold long-duration nominal sovereign bonds as the primary defensive allocation” — that is a claim we can test. Many celebrated investment pronouncements dissolve at this gate because they are too vague to implement or too context-specific to transport.
We also distinguish what an author actually said from what their popularizers claim they said. Buffett’s instruction to put 90% in an S&P 500 index fund and 10% in short-term Treasuries was for cash left in trust for his wife, not a universal allocation. Graham’s 25–75% stock/bond range was for a U.S. defensive investor in 1973, with instruments and yields specific to that moment. If we cannot pin the claim to a primary source, context, and intended audience, we cannot transport it.
Gate 2: Mechanism
Second: why should it work — economically or behaviourally?
“Because Graham said so” is not a mechanism. “Because it backtested well” is not a mechanism. A mechanism is a causal account: costs compound, so lower costs leave more return. Stock returns are skewed, so broad ownership reduces the chance of missing the few extreme winners. Duration amplifies sensitivity to discount-rate changes, so long bonds gain when yields fall and lose when yields rise.
Without a mechanism, we have no way to judge whether a historical result was luck or structure, and no way to assess whether it will persist when conditions change.
Gate 3: Evidence and counterevidence
Third: what evidence, transparent and reproducible, supports the mechanism? And what evidence challenges it?
Evidence hierarchy:
- Primary writings and official sources: What the author actually said, in context. Official documentation of instrument mechanics (e.g., TreasuryDirect on TIPS, index-provider methodology).
- Peer-reviewed academic research: The gold standard, but with important caveats about sample period, geography, investability, and look-ahead bias.
- Practitioner research with transparent methodology: Vanguard, S&P, major asset managers — useful when the method is disclosed and the limitations acknowledged.
- Raw retrievals and market-data probes: Leads for investigation, not conclusions.
- Anecdotes, narratives, and unsourced charts: Not evidence.
Crucially, we require counterevidence. A claim that ignores the studies that challenge it is propaganda, not analysis. For every rule, we ask: what is the strongest argument that could genuinely change this?
Gate 4: Boundary and failure mode
Fourth: when does this rule fail, and under what conditions?
No rule is universal. The question is whether the boundaries are known and acknowledged:
- Horizon: Does the rule require a long holding period? What happens if the investor must liquidate early?
- Currency: Does the rule assume the investor’s liabilities are in the same currency as the assets?
- Regime: Does the rule depend on a particular inflation, growth, or covariance environment?
- Construction: Is the rule sensitive to how the asset is packaged (individual bond vs. fund, physical vs. futures, etc.)?
- Behaviour: Can a real human, experiencing real fear and greed, follow this rule through a drawdown?
A rule whose boundaries are unknown is more dangerous than a rule we know is bad.
Gate 5: Modern relevance
Fifth: has market structure changed in a way that alters the mechanism, implementation, or failure boundary?
Some things change. Low-cost global index funds were not available to Graham in 1973; they are now. Inflation-linked bonds did not exist in major markets before 1981; they now do. The post-1980 disinflation that powered bond returns is one regime, not the permanent baseline.
Other things do not change: costs compound, diversification cannot eliminate systematic risk, duration is sensitivity to discount-rate changes, and leverage can force ruin.
The test is specific: a modern development changes a rule only if it alters the mechanism, the investable implementation, the relevant liability or currency, or the probability/severity of a failure mode in a way supported by more than a current narrative.
“The market feels different now” is not sufficient.
Gate 6: Disposition
Sixth: what is the final classification?
| Disposition | Definition |
|---|---|
| Durable core | Broad mechanism, robust across plausible regimes, implementable without precise forecasts. |
| Conditional tool | Valid only for a named job, condition, construction, or investor input. |
| Unsupported/overstated | The proposed action outruns its mechanism or evidence, conceals construction, or is a forecast presented as a rule. |
A claim can be partially durable: risk transparency is durable; a specific risk-parity allocation is conditional. Ruin avoidance is durable; a 90/10 barbell is underspecified.
The evidence hierarchy
Not all sources are equal. The framework uses an explicit hierarchy:
| Tier | Source type | Examples | Status |
|---|---|---|---|
| 1 | Primary writings and official sources | Graham’s The Intelligent Investor (1973), Sharpe (1991), TreasuryDirect on TIPS mechanics, index-provider methodology documents | Evidence of what was said and how instruments work |
| 2 | Peer-reviewed academic research | Bessembinder (2018), Odean (1998), French–Poterba (1991), Campbell–Viceira, Frazzini–Israel–Moskowitz | The gold standard, but subject to sample, geography, investability, and look-ahead limits |
| 3 | Practitioner research with transparent methodology | Vanguard rebalancing studies, S&P factor attribution, BIS working papers | Useful when method is disclosed and limitations acknowledged |
| 4 | Raw retrievals and market-data probes | Archived research streams, pre-gate price data | Leads for investigation; not conclusions |
| 5 | Anecdotes, narratives, and unsourced claims | “Everyone knows bonds are safe,” “Gold always hedges inflation” | Not evidence |
A Tier 3 source can inform a rule but cannot override a Tier 2 source with a contradictory finding. A Tier 5 source cannot inform anything.
What counts as “evidence” — and what doesn’t
Six controls govern the use of evidence in this project:
- No authority worship. A name or a quotation is a lead, not proof. The framework’s core constraints (K1–K4) are supported by independent mechanism and evidence — Sharpe (1991) for cost arithmetic, Bessembinder (2018) for diversification, Kelly/ergodicity framework for survival, instrument mechanics for liquidity. None depends on “because Graham/Buffett/Bogle said so.”
- No backtest optimization. Selecting the allocation that performed best over one historical period guarantees nothing about the next period. DeMiguel et al. checked a narrower claim about optimization error; the general point — that optimized historical weights are fragile — is well-established.
- No current-observation trading. “Yields are high,” “concentration is extreme,” “inflation is rising” describe the present. A durable rule must survive changes in these variables. A current observation is not a durable mechanism.
- No one-variable macro stories. “Debt/GDP is high, therefore inflation” omits maturity, currency, holder base, primary balance, r – g differential, monetary regime credibility, external position, and institutional context. Sargent–Wallace and Leeper provide the theoretical framework; the one-variable claim collapses under it.
- Distinguish ex-ante from hindsight. An indicator available only after the fact is not a decision rule. A retrospective label based on the asset return being explained is not admissible as a state definition.
- Flag sample limits. U.S.-only, short-sample (crypto’s ~15 years), pre-cost, pre-tax, and survivorship-biased evidence must be labeled as such. The post-1980 disinflation bias affecting most 60/40 and safe-withdrawal evidence must be explicitly acknowledged.
The empirical escalation rule
Data work is done only when a result could change a rule’s disposition or a material boundary. It is not done because data is available or because a backtest would be interesting.
Before any empirical test, we specify: the rule being tested, the genuine rival hypothesis, the exact series and vehicle, total-return versus price status, currency, sample limits, transformation, costs, and interpretation limits. We do not optimize weights, select dates after seeing returns, or infer causality from a chart.
This is a high bar. It means many questions remain unresolved in the framework — and that is a feature, not a bug. Pretending we know what we don’t know is the most expensive error in investing.
Claims that failed at each gate
To make the method concrete, here are examples of widely-circulated claims that failed at specific gates — and why.
Failed at Gate 1 (imprecise or context-specific claim)
“Buffett recommends 90% S&P 500, so that is the right allocation.” The claim is specific enough to implement, but it fails the context test: Buffett’s instruction was for cash left in trust for his wife, not a universal global allocation. The same passage says his wealth is in Berkshire Hathaway stock and the 90/10 applies to the remaining cash portion. Transporting it without the context is an authority-extrapolation error.
Failed at Gate 2 (no mechanism)
“Debt/GDP is above X%, therefore inflation is inevitable.” The claim proposes an allocation action (reduce nominal bonds) but provides no mechanism connecting a debt ratio to a price-level outcome. The missing causal chain: deficits → monetary accommodation → credit expansion → demand exceeding supply → price increases. Each link requires conditions (maturity structure, currency of issuance, holder base, central bank reaction function, output gap) that the one-variable claim ignores. Sargent–Wallace and Leeper show that fiscal-monetary outcomes depend on the interaction of policy regimes, not a single ratio.
Failed at Gate 3 (counterevidence stronger than evidence)
“Gold reliably hedges CPI inflation at practical portfolio horizons.” The evidence for this claim relies on very-long-horizon (century-scale) anecdotes. The counterevidence — Erb and Harvey’s finding that the gold/CPI ratio has ranged from ~1:1 to over 8:1 historically — shows that at horizons relevant to portfolio construction, the relationship is unreliable. Gold can fall during inflationary periods and rise during disinflationary ones. The mechanism is overwhelmed by real-rate, currency, and sentiment effects. (Disposition: U11 — overstated; gold’s CPI-hedge job is unsupported.)
Failed at Gate 4 (unknown boundaries and failure modes)
“Long nominal bonds are always safe and always hedge equities.” This claim failed in 2022, but the deeper problem was that it was never bounded. Stock–bond covariance is not a fixed property — it depends on whether growth/inflation shocks dominate. Campbell, Sunderam, and Viceira document sign changes. The BIS and ECB link the post-2021 positive correlation to the inflation environment. The claim “bonds always hedge equities” ignored the failure mode (inflation/real-yield shock) and the regime dependence of the correlation. (Disposition: U6 — unsupported universal hedge claim.)
Failed at Gate 5 (modern relevance ignored or overclaimed)
Two symmetric errors: ignoring genuine structural change, and overclaiming it.
Error A — ignoring change: Treating the post-1980 disinflation (CPI from ~15% to ~2%, 10-year yields from ~16% to <1%) as the permanent baseline. Substantial safe-withdrawal and 60/40 evidence draws on this uniquely favourable sample. Pfau (2010), applying the same method to 17 developed markets over 1900–2008, finds substantially lower sustainable rates — a direct refutation of the claim that the post-1980 sample is representative.
Error B — overclaiming change: “QE/reserves/M2 growth mechanically predicts CPI inflation.” These are distinct balance-sheet and monetary objects. Transmission depends on credit, fiscal transfers, money demand, reserve remuneration, supply constraints, and expectations. The change in monetary operations (ample-reserve floor systems) is an implementation detail, not a changed portfolio mechanism. The claim confuses a visible change in central bank practice with a mechanical inflation rule. (Disposition: U4 — unsupported mechanical/timing rule.)
Failed at Gate 6 (disposition overreach)
The Permanent Portfolio’s 4×25 allocation. The mechanism (scenario-based diversification) is durable. But the specific weights (25% each) overreach: equal capital is not equal risk, 25% gold is the largest active bet, 25% long nominal duration creates substantial inflation sensitivity, and the regime separation fails under stagflation. The durable ideas are retained via the diversification constraint, job definition, precommitment, and simplicity disciplines; the specific allocation is not adopted. The error was claiming a complete, all-weather portfolio when the construction was conditional on a specific economic-regime mapping that does not hold under all states.
The red-team discipline
Every core rule was tested against the strongest counterargument that could genuinely change it. Not a straw man — the real thing. This internal adversarial process found no contradiction sufficient to reject the framework, but it also did not prove the rules are minimal, complete, or robust for every investor.
The red team found:
- That some rules needed boundary clarifications (e.g., cost discipline requires implementation diligence beyond the headline expense ratio; liquidity separation has an inflation-erosion tension).
- That unsupported generic numeric ranges and automatic review triggers should be withdrawn — and they were.
- That the framework’s core constraints survive not because they are clever, but because they are narrow: they constrain the portfolio design space without dictating asset weights.
Key idea: A disciplined method is the only defense against the human tendency to confuse a good story with a good reason. The chain — claim, mechanism, evidence, boundary, modern relevance, disposition — is slow, but the alternatives (authority, backtest, narrative) have worse track records.