Portfolio Optimization
Portfolio Optimization
The last post gave us alpha: the gap between an asset’s expected return and what its risk exposure warrants. As a pricing measure, alpha marks an asset as cheap or expensive relative to its systematic risk. It does not tell us how far a portfolio should tilt toward an asset whose alpha is positive.
Answering that takes more machinery than we have. The variance formula we have so far covers two assets, and real portfolios hold hundreds of securities, more than that formula can handle by hand. This post scales the risk arithmetic to assets, states the optimization problem built on it, and then puts alpha to work choosing weights.
The Covariance Matrix
For two assets A and B held at weights and , post 8 gave the portfolio variance as:
That expression has 3 unique terms: two variances and one covariance. With 3 assets we get 6, with 10 we get 55, and with 500 we get 125,250. The count grows as .
For assets, the general form is a double sum over every pair:
When , and the term is asset ’s own variance. When , the term is the cross-covariance between assets and .
Matrix notation compresses the whole double sum into three symbols:
where is the weight vector and is the covariance matrix. The diagonal entries are variances () and the off-diagonal entries are covariances (). The matrix is symmetric, since .
We write the product out for and recover the two-asset formula:
Multiplying through gives , and two substitutions recover the form above. The first is cosmetic, since is . For the second we borrow post 8’s definition of correlation, covariance divided by both standard deviations, and rearrange it into to replace the cross-term.
For 3 assets with , , and correlations , , , each diagonal entry is and each off-diagonal entry is . In percent squared, the covariance matrix is:
The first diagonal entry is , and the entry linking the first two assets is . The matrix has three variance terms on the diagonal and three unique covariance terms off it. At we can still write the matrix out by hand. At we cannot, which is what the notation is for.
Drag the slider below to add assets to the matrix and watch the term counts move.
The highlighted cells along the diagonal are variance terms and the rest are covariances, with the faded cells above the diagonal repeating the ones below.
Why Covariances Dominate
For assets, the covariance matrix has variance terms (the diagonal) and unique covariance terms (the lower triangle). The ratio of covariance terms to variance terms is to , and the two counts diverge as grows:
| Assets | Variance terms | Covariance terms | Total unique |
|---|---|---|---|
| 2 | 2 | 1 | 3 |
| 10 | 10 | 45 | 55 |
| 100 | 100 | 4,950 | 5,050 |
| 500 | 500 | 124,750 | 125,250 |
At , covariance terms outnumber variance terms 249.5 to 1. Portfolio risk is overwhelmingly determined by how assets move together, not by how volatile they are individually. This is the algebraic version of the diversification insight we reached in post 8: individual risk washes out in a large portfolio, but correlated movements survive. When we add a new security to a portfolio, its own variance matters less than how it covaries with everything already held.
Mean-Variance Optimization
An investor wants the highest expected return available at a chosen level of risk, or equivalently the lowest risk at a target expected return. Mean-variance optimization states that goal as a constrained minimization:
subject to:
where is the vector of expected excess returns (each asset’s ), is the target portfolio excess return, and is a vector of ones enforcing that the weights sum to one.
The objective is to make portfolio variance as small as the constraints allow. The first constraint pins expected return to the target, and the second keeps the portfolio fully invested. We solve it at every possible and trace out the efficient frontier from post 9, now for assets instead of two.
With a risk-free asset available, the problem simplifies to a search for the highest Sharpe ratio:
The solution is the tangent portfolio. For an investor with risk aversion , the optimal risky portfolio is:
The optimal weights are the inverse covariance matrix times the vector of expected excess returns, scaled by . The covariance matrix acts as a penalty for redundant risk, down-weighting assets that covary heavily with the rest of the portfolio, while assets with high expected returns get more weight. The risk aversion parameter scales the whole position, so a less risk-averse investor ( small) takes larger positions and a more risk-averse investor ( large) takes smaller ones.
For the tangent portfolio itself, does not matter. Every value of produces the same portfolio once we normalize the weights to sum to one, only at a different scale. This is Two Fund Separation again: decides how much of the tangent portfolio an investor holds versus the risk-free asset, and the tangent portfolio is the same for everyone. Normalizing gives us the fully invested tangent weights directly, with no risk-free position:
The denominator is a scalar, the sum of the unnormalized weights. Dividing by it rescales the vector so that .
This is Markowitz’s result in its general form. The efficient frontier, Two Fund Separation, and the tangent portfolio all fall out of this one optimization.
The hard part in practice is not the math. It is estimating and reliably. Small errors in expected returns produce large swings in the optimal weights, so unconstrained mean-variance optimization is notoriously sensitive to its inputs. The unconstrained solution can also produce negative weights (short positions), which many investors cannot take. Adding a no-shorting constraint () breaks the closed-form solution and calls for numerical optimization, typically quadratic programming.
Alpha Relative to Any Portfolio
In post 10 we measured alpha against the market, under the CAPM, or against a factor model. The definition generalizes to any reference portfolio. For asset measured against a portfolio :
where . The formula is post 10’s, with only the reference portfolio changed. CAPM alpha uses the market, and nothing stops us from using the portfolio we already hold, or any other benchmark. The question shifts from “does this asset beat the market?” to “does this asset offer more return than its risk contribution to my portfolio warrants?”
Alpha as a Gradient
The marginal improvement to a portfolio’s Sharpe ratio from adding a small amount of asset is proportional to that asset’s alpha:
In words: the rate at which the Sharpe ratio improves as we begin tilting toward asset equals the asset’s alpha, measured against the portfolio held now, divided by that portfolio’s standard deviation.
Alpha is excess return after adjusting for risk, so an asset offering more return than its risk contribution warrants improves the risk-return tradeoff when we fold it in. A larger alpha makes the improvement steeper. A more volatile starting portfolio makes it shallower, since is in the denominator.
That result reframes alpha from a scorecard for past performance into a direction. Positive alpha points toward a better portfolio, negative alpha points away from one, and zero alpha says the allocation to that asset is already right.
Suppose we hold a portfolio with a Sharpe ratio of 0.53 (, , ), and an asset outside it offers , , and correlation with the portfolio. We compute the asset’s beta relative to the portfolio, , and its alpha, . The gradient is , so each percentage point of tilt initially raises the Sharpe ratio by about 0.0027.
Tilting Toward Alpha
Suppose asset has against the portfolio we hold. Tilting toward it raises the Sharpe ratio, and two things change as we tilt. The portfolio itself changes, since it now holds some of asset . The asset’s beta against the new blended portfolio changes with it, because the composition has shifted. The alpha shrinks.
At some tilt the alpha of the asset against the blended portfolio reaches zero, and no further tilt improves anything. Past that point the alpha turns negative and the Sharpe ratio declines with it.
That endpoint gives us an optimality condition. The maximum-Sharpe portfolio is the one where every asset has zero alpha relative to it, its return matching what its risk contribution warrants.
Drag the tilt slider below and watch the Sharpe ratio rise, peak, and fall.
At the default settings the asset has positive alpha against the portfolio, and the dashed line marks the tilt where the Sharpe ratio peaks. Raise the asset’s expected return and the curve lifts at every positive tilt.
CAPM as an Optimality Condition
The CAPM entered in post 10 as a pricing model. Read in the language of this post, it is a statement about optimality. The CAPM assumes the market portfolio has the highest Sharpe ratio among all portfolios of risky assets, and if that holds, every asset must have zero alpha relative to it. The CAPM equation
is the condition for every asset , with the market portfolio as the reference. Beyond telling us what returns should be, the model says the market portfolio is mean-variance optimal. No tilt away from market weights can improve the Sharpe ratio.
Pricing and portfolio construction are the same statement read in two directions. Grant that the market portfolio is optimal and we get CAPM pricing for free. Claim instead that some asset is mispriced and has positive alpha, and we have implicitly said the market portfolio is suboptimal. The alpha-tilt result then shows how to improve on it.
What’s Next
We now have a complete toolkit for building portfolios and improving them. We have not yet asked whether any of this analysis yields an edge. If the market already reflects all available information, alpha should be zero everywhere and the tangent portfolio should be the index. The final post takes up the efficient market hypothesis and what it means for everything we’ve built.