Email usBook a call (opens in a new tab)
Quant

Sharpe vs Sortino Ratio: Why One Lookback Is Never Enough

Sharpe and Sortino are simple to define and easy to get wrong. Correct formulas, honest annualisation, and why a single lookback window hides more than it shows.

Every strategy report, fund factsheet and trading dashboard quotes a Sharpe ratio. Many quote a Sortino ratio next to it. Few say how either was calculated, over what period, at what frequency, or against what risk-free rate. Those choices can move the number by more than the difference between a good strategy and a mediocre one. This piece sets out the correct definitions, the annualisation rules, the downside deviation detail that most implementations get wrong, and the reason we never publish a single-window ratio on its own.

The definitions

Sharpe ratio

The Sharpe ratio is the mean excess return divided by the standard deviation of excess returns:

Sharpe = mean(R - Rf) / stdev(R - Rf)

R is the strategy return per period and Rf is the risk-free return for the same period. Use the sample standard deviation (ddof=1). The result is a per-period ratio until you annualise it.

Strictly, Sharpe's 1994 revision defines the denominator as the standard deviation of the excess return series. When the risk-free rate is close to constant over the window, this is almost identical to the standard deviation of raw returns, but computing it on the excess series is the correct form and costs nothing.

Sortino ratio

The Sortino ratio replaces total volatility with downside deviation, which penalises only returns below a target:

Sortino = mean(R - T) / DD

DD = sqrt( (1 / N) * sum( min(0, R_i - T)^2 ) )

T is the target or minimum acceptable return per period. It is often zero or the risk-free rate; state which. N is the total number of observations.

That last point matters. The sum runs over all observations, with the positive excess returns contributing zero, and the denominator is N, the full sample size. A common bug is to filter to the negative returns first and then take their standard deviation. That calculation does two wrong things: it divides by the count of negative periods rather than all periods, and it subtracts the mean of the negative returns rather than measuring distance from the target. The result understates downside risk for strategies with few losing periods and inflates their Sortino, sometimes dramatically.

How they compare

Property Sharpe ratio Sortino ratio
Risk measure Standard deviation of excess returns Downside deviation below target
Penalises upside volatility Yes No
Needs a target return Risk-free rate Target T (state it)
Sensitive to return asymmetry No, treats both tails equally Yes, rewards positive skew
Typical relationship Baseline About 1.4 times Sharpe for symmetric returns near zero mean
Stability on small samples Moderate Lower, few downside observations drive it

The rough 1.4 multiple is worth knowing as a sanity check. For a symmetric distribution with mean close to the target, downside deviation is about the standard deviation divided by the square root of two, so Sortino lands near 1.41 * Sharpe. If your Sortino is five times your Sharpe, either the return distribution is strongly skewed (worth investigating) or the downside deviation is computed incorrectly (more likely).

Annualisation

Ratios are compared on an annual basis. To annualise a per-period ratio, multiply by the square root of the number of periods per year:

annualised ratio = per-period ratio * sqrt(P)

This follows because the mean scales with P and the standard deviation (or downside deviation) scales with sqrt(P), assuming returns are independent and identically distributed.

The right P depends on what a row in your data represents:

  • Daily returns on exchange trading days: P = 252 is the convention for equities and futures. Use it only if your series contains trading days only.
  • Daily returns including weekends: crypto and some FX-style series trade every day. Use P = 365. Using 252 on a 365-day series understates the annualised ratio by a factor of sqrt(252/365), about 0.83.
  • Weekly: P = 52. Monthly: P = 12.

Do not mix. If one strategy's returns are daily and another's are monthly, resample both to the same frequency before comparing. Annualised Sharpe ratios computed from different frequencies are not comparable when returns are autocorrelated, which trading strategies often are.

The independence assumption

The sqrt(P) rule assumes no serial correlation. Andrew Lo showed in 2002 that positive autocorrelation (common in strategies holding illiquid positions or smoothing marks) makes the square-root rule overstate the annual Sharpe, while negative autocorrelation (common in mean-reversion) makes it understate. For a serious evaluation, check the autocorrelation of the return series, and if it is material, apply an adjusted scaling or report the ratio at the native frequency alongside the annualised figure.

Handling the risk-free rate

Three rules cover most cases.

  1. Match the period. Convert the annual risk-free rate to the return frequency before subtracting. The geometric conversion is (1 + rf_annual) ^ (1 / P) - 1. Dividing by P is a close approximation for low rates and short periods.
  2. Use a time series where rates moved. Between 2021 and 2023 short-term USD and GBP rates rose from near zero to over 5%. A constant rate across that period distorts excess returns in both halves. Use daily SOFR, SONIA or a T-bill series aligned to your return dates.
  3. Be explicit. Some reports use zero. That is acceptable if declared, but a zero-rate Sharpe is not comparable with a rate-adjusted one. At 5% rates, a strategy returning 8% at 10% volatility has a Sharpe of 0.8 with zero Rf and 0.3 with the correct one.

Why one lookback is never enough

A single Sharpe ratio is an estimate with a standard error, not a fact about the strategy. For independent, normally distributed returns, the standard error of a per-period Sharpe is approximately:

SE(Sharpe) ≈ sqrt( (1 + 0.5 * Sharpe^2) / N )

with N the number of observations. On 252 daily observations and a modest true Sharpe, the annualised standard error is roughly 1.0. A one-year Sharpe of 1.5 is therefore consistent with a true value anywhere from about 0.5 to 2.5. Over 30 trading days it is close to meaningless.

To show how loud that noise is, we simulated 1,300 trading days of returns from a single fixed process: normally distributed, daily mean 0.04%, daily volatility 1%, 4% annual risk-free rate. The true annualised Sharpe is about 0.39. Then we computed both ratios across six trailing windows:

Window Observations Sharpe Sortino
30 days 22 2.78 4.95
90 days 65 -0.78 -1.14
180 days 130 -1.22 -1.63
1 year 261 0.74 1.11
3 years 783 0.92 1.34
Full history 1,300 0.41 0.58

Nothing changed in the process. The strategy did not get better or worse. Yet any one of those rows, quoted alone, tells a different story: a star in the last month, a failure over six months, solid over three years. Real strategies add genuine regime dependence on top of this sampling noise, which makes a single window even less trustworthy.

Showing the full ladder of lookbacks does three things:

  • Reveals regime dependence. A trend-following system may show a high Sharpe over a volatile year and a negative one over a quiet quarter. That pattern is information about the strategy, not noise to average away.
  • Exposes cherry-picking. If the headline window happens to be the best one, the ladder makes that obvious.
  • Separates signal from sample size. Short windows with wildly different values next to long windows with stable values tell you the short windows are mostly noise.

We present 30D, 90D, 180D, 1Y, 3Y and full-history figures side by side, with the observation count on every row.

Common errors

Mixed frequencies

Calculating the mean from daily data and the volatility from monthly data, or annualising monthly returns with sqrt(252), produces numbers that look plausible and are wrong. Resample once, compute once.

Survivorship bias

If you evaluate a universe of strategies, funds or symbols, include those that were stopped, closed or delisted. Ratios computed only on survivors are biased upward, because the failures have been removed from the sample.

Look-ahead bias

The return in row t must only use information available at t. Signals computed on close prices and filled at that same close, or position sizing that uses full-sample volatility, leak future information into the backtest. The Sharpe of a leaky backtest is not a pessimistic or optimistic estimate; it is a measurement of a different, impossible strategy.

Small samples

Below a few hundred observations, Sortino is particularly unstable because the downside deviation may rest on a handful of losing days. Report N with every ratio, and treat ratios on fewer than about 60 observations as descriptive only.

Returns on the wrong base

Compute returns on account equity (or allocated capital) including costs, financing and fees, not on notional or on gross P&L. A strategy using leverage reports a very different volatility depending on the base.

A clean pandas implementation

The code below computes both ratios correctly and builds the lookback ladder. It expects a pandas Series of simple periodic returns indexed by date.

import numpy as np
import pandas as pd


def per_period_rf(annual_rf, periods_per_year):
    """Convert an annual risk-free rate to a per-period rate (geometric)."""
    return (1 + annual_rf) ** (1 / periods_per_year) - 1


def sharpe_ratio(returns, annual_rf=0.0, periods_per_year=252):
    r = returns.dropna()
    excess = r - per_period_rf(annual_rf, periods_per_year)
    sd = excess.std(ddof=1)
    if len(excess) < 2 or sd == 0:
        return np.nan
    return excess.mean() / sd * np.sqrt(periods_per_year)


def sortino_ratio(returns, annual_target=0.0, periods_per_year=252):
    r = returns.dropna()
    excess = r - per_period_rf(annual_target, periods_per_year)
    downside = np.minimum(excess, 0.0)
    dd = np.sqrt((downside ** 2).mean())   # divide by ALL observations
    if len(excess) < 2 or dd == 0:
        return np.nan
    return excess.mean() / dd * np.sqrt(periods_per_year)


def lookback_table(returns, annual_rf=0.0, periods_per_year=252):
    windows = {"30D": 30, "90D": 90, "180D": 180, "1Y": 365, "3Y": 3 * 365}
    end = returns.index.max()
    rows = []
    for label, days in windows.items():
        window = returns[returns.index > end - pd.Timedelta(days=days)]
        rows.append((label, len(window),
                     sharpe_ratio(window, annual_rf, periods_per_year),
                     sortino_ratio(window, annual_rf, periods_per_year)))
    rows.append(("Full", len(returns),
                 sharpe_ratio(returns, annual_rf, periods_per_year),
                 sortino_ratio(returns, annual_rf, periods_per_year)))
    return pd.DataFrame(
        rows, columns=["window", "n_obs", "sharpe", "sortino"]
    ).set_index("window")

A few design notes:

  • Windows are defined in calendar days and the observation count is reported, so a 30-day window on trading-day data correctly contains about 21 or 22 rows.
  • np.minimum(excess, 0.0) keeps every observation, so the mean of the squared values divides by N, as the definition requires.
  • The constant risk-free rate is a simplification for clarity. In production, pass a Series of per-period rates aligned to the return index and subtract element-wise.
  • Zero-variance and tiny samples return NaN rather than infinity or a misleading number.

Key takeaways

  • Sharpe uses total volatility of excess returns. Sortino uses downside deviation below a stated target.
  • Downside deviation divides by all observations, not just losing ones. Filtering first inflates Sortino.
  • Annualise with sqrt(P), where P matches the actual rows in your data: 252 for trading days, 365 for every-day markets.
  • Convert the risk-free rate to the return frequency and use a time series when rates have moved.
  • A single window is one noisy sample. Show 30D through full history with observation counts.
  • Check for mixed frequencies, survivorship, look-ahead and small samples before trusting any ratio.

How we apply this

Our analytics code treats performance metrics as calculations with inputs that must be declared: frequency, periods per year, risk-free source, target return and window. Those parameters are stored next to every figure, so a number on a dashboard can always be traced back to how it was produced. The lookback ladder is the default view, not an optional extra.

We run this on our own live and research systems, which keeps us honest about how much a short window can mislead. If you need performance reporting, backtest evaluation or investor-facing analytics built to that standard, see our quantitative analytics services and fund technology work.


This article is for information and education only and is not investment advice. See our risk disclaimer.

Start a project

Need this engineered, not just explained?

We turn research like this into production software for trading and investment businesses.

Prefer email or phone? hello@aurionlabs.io · +44 7832 617626