The Efficient Frontier, Out of Sample

The frontier is drawn from a sample of returns, and the portfolio on its best point was picked because it looked best in that sample. Hold it through the years that follow and see what survives.

Loading interactive simulation...

An optimiser cannot tell a risk premium from a lucky sample 🖖

Fit these five asset classes on 1928–1957 and the best-Sharpe portfolio puts 93.4% into 10-year Treasury bonds. That is not a judgement about bonds; it is arithmetic. Bills paid 1.19% a year over that stretch and bonds beat them at a ninth of the volatility of equities, so bonds carried the highest reward per unit of risk in the sample, and finding exactly that is the optimiser's whole job. Held through 1958–1987 β€” three decades of rising inflation and rising rates β€” it scored 0.029 against the 0.700 it had promised, while equal weights scored 0.484. Its raw return actually rose, from 3.50% to 6.36%, which is why the raw return is the wrong thing to look at: bills rose faster still, to 6.10%, so a portfolio that had promised 2.31% a year above cash delivered 0.25% while its volatility nearly tripled. The optimiser was not wrong about the past. It was asked a question about the future and answered a question about the sample, and nothing in the method distinguishes the two.

The three portfolios finish in exactly reverse order 🖖

On the window the page opens with, the table promises a Sharpe ratio of 0.829 for the optimised tangency portfolio, 0.777 for minimum variance and 0.565 for equal weights. Held through the following 16 years they deliver 0.494, 0.524 and 0.757: the same three portfolios, in reverse. The order is not luck, it is how much of the sample each one was fitted to β€” tangency uses the estimated means and the estimated covariances, minimum variance uses only the covariances, equal weights uses neither. Averaged over all 59 twenty-year fit-and-hold windows the ranking is the same: 0.297, 0.460, 0.558.

The portfolio built to be calm was the wild one 🖖

Minimum variance promised 5.32% volatility on this window β€” lower than the tangency portfolio’s 5.67% β€” and delivered 8.60%, higher than tangency’s 7.93%. Lowest promise, highest outcome. It is not one window: sweep every fit length from 10 to 30 years against every hold length from 5 to 30 and the tangency portfolio’s realised volatility comes out above its fitted volatility in 22,203 of 33,579 of them. An optimiser hunts for the corner of the estimated covariance matrix where risk looks smallest, and the corners where an estimate is smallest are exactly the ones where it is most likely to be too small.

Problem solved in full

  1. The best-Sharpe mix of US stocks and 10-year Treasuries 7 steps

    Two assets, one window: US stocks and 10-year Treasuries fitted on 1928–1976. Derive the best-Sharpe mix by hand, check it against the panel — then work out how many of its printed digits the data can actually support.

    1. Click Stocks and bonds only to fit the frontier on this window. Six numbers go in and nothing else: each sleeve’s mean annual return, each sleeve’s volatility, their correlation, and the bill rate that defines risk-free.

    2. Sharpe ratios are measured above cash, so subtract the bill from both means. Stocks earned about eleven times the bond sleeve’s excess. The covariance is slightly negative — a small number that does real work two steps below.

    3. A mix holding w in stocks earns the weighted average of the two excesses. Its volatility is not a weighted average: the cross term is the whole of diversification, and with a negative covariance it subtracts.

    4. Before trusting the formula, test it where the answer is already on the page. Put w = ½ and the tool’s own Equal weights (1/N) row reads 4.52% above cash at 11.30% volatility. Both match, so the six inputs are the right six.

    5. Now maximise. Because the Sharpe ratio is a ratio, its maximum need not sit at either end — and with a negative covariance it does not. Setting the derivative to zero clears the square root and leaves a formula in the same six inputs.

    6. Substitute. Open The weights each method chose and the tangency row reads 27.2% stocks, 72.8% bonds; the table above it reads 6.72% predicted volatility and a predicted Sharpe of 0.418. A portfolio a quarter in stocks is 4.6 points quieter than the 50/50 mix.

    7. Now price the inputs, which the tool never does. μs is not a fact about the world; it is the average of 49 annual returns whose spread is 22.351%, so it carries a standard error of its own. Put μs one standard error either side and re-solve the same formula.

    Answer

    27.2% stocks, 72.8% bonds — printed to a tenth of a point on a quantity the data pins down to about ±7 points. One standard error on the equity mean alone swings the optimal weight from 19.3% to 33.4%, and that is with the two volatilities, the correlation and the bill rate all still treated as known exactly. The objective is no help in narrowing it: evaluate S(w) at 24.1% and at 30.9% and both round to 0.417, so the panel’s own three decimals cannot separate weights seven points apart. That is the frontier’s quiet joke — its most confident-looking output is its least determined one. Which is why equal weights, estimating nothing at all, delivered 0.551 here against the optimiser’s 0.440.

Learning path

Risk, measured

References (6)

Example problems

  • Promise vs delivery - Fitted on 1990-2009 and held through 2010-2025: the best-Sharpe portfolio's score falls from 0.829 to 0.494 while equal weights rise from 0.565 to 0.757.
  • All in on bonds - Fitted on 1928-1957, the optimiser puts 93.4% into Treasury bonds - and scores 0.029 across the three decades of rising rates that followed.
  • When it worked - Fitted on 1970-1994 and held through 1995-2009, the optimised portfolio does beat equal weights: 0.772 against 0.658.
  • Stocks and bonds only - Two assets and a 49-year fit, where the frontier is a simple curve and the tangency portfolio holds 27.2% equities.