Small Group Tutorials

Here to help students catch up, keep up, and move ahead. Book a consultation here.

How Nelson–Siegel–Svensson Yield-Curve Algorithms Fit Spot and Forward Rates: Level/Slope/Curvature Loadings, Nonlinear Calibration, Identifiability, Arbitrage and Extrapolation Failure

Reader question: Market bonds exist only at scattered maturities and contain coupon, liquidity and pricing noise. How can a small set of parameters turn those observations into a smooth spot-rate and forward-rate curve without fitting every maturity independently?

Nelson–Siegel and Nelson–Siegel–Svensson (NSS) algorithms use a compact family of exponential basis functions. Their coefficients produce long-run level, short-end slope and one or two medium-term curvature components. Calibration then chooses the parameters that minimize pricing or yield errors across the observed instruments.

This article owns one precise computational job: cross-sectional parametric smoothing of a yield curve with Nelson–Siegel/Svensson basis functions and nonlinear decay-parameter calibration. It does not own instrument bootstrapping, historical yield-curve PCA, HJM risk-neutral dynamics or a guarantee of dynamic no-arbitrage.

This is public mathematical and computational education. It is not a bond recommendation, a forecast of future interest rates or a claim that the fitted long-run parameter is a true long-run economic equilibrium rate.

1. Why a parametric yield curve exists at all

A market may contain many coupon bonds, bills and swaps, but not one liquid zero-coupon instrument at every maturity.

A smooth curve is useful for:

  • interpolating yields between observed maturities;
  • producing zero-coupon discount factors;
  • deriving instantaneous and discrete forward rates;
  • comparing securities against a common term structure;
  • publishing stable reference curves;
  • compressing a cross-section into a few interpretable parameters.

The challenge is balancing fit and smoothness. A curve flexible enough to pass through every noisy quote can create implausible oscillating forward rates; a curve that is too rigid can miss economically important shape.

2. Nelson–Siegel starts from a forward-rate shape

One common Nelson–Siegel instantaneous forward-rate representation is:

f(τ)=β0 + β1e−τ/τ1 + β2(τ/τ1)e−τ/τ1.

Here τ is maturity in years and τ1>0 is a decay parameter.

Integrating the forward curve gives the continuously compounded zero yield:

y(τ)=β0 + β1L1(τ) + β2L2(τ),

with:

L1(τ)= [1−e−τ/τ1]/(τ/τ1)

and:

L2(τ)=L1(τ)−e−τ/τ1.

3. Long-maturity limit: β0

As τ→∞:

L1(τ)→0,

L2(τ)→0.

Therefore:

y(∞)=β0.

This is why β0 is commonly called the level parameter.

But it is a model asymptote. If the available data stop at 20 or 30 years, β0 is partly an extrapolation parameter rather than a directly observed infinite-maturity yield.

4. Short-maturity limit: β0+β1

As τ→0:

L1(τ)→1,

L2(τ)→0.

So:

y(0)=β0+β1.

β1 therefore controls the difference between the short end and the long-run level and is commonly interpreted as a slope component.

5. The curvature loading

L2(τ) starts near zero, rises to a hump and returns toward zero at long maturity.

β2 therefore changes the middle of the curve more than the extreme short or long ends.

τ1 controls where that hump occurs and how quickly the slope loading decays.

This produces the familiar level/slope/curvature interpretation without requiring the statistical eigenvectors used in PCA.

6. Svensson adds a second curvature term

The Svensson extension adds:

β3L3(τ),

where:

L3(τ)= [1−e−τ/τ2]/(τ/τ2) − e−τ/τ2

and τ2>0 is a second decay parameter.

The full NSS zero-yield curve is:

y(τ)=β0+β1L1(τ)+β2L2(τ)+β3L3(τ).

The second hump allows the curve to represent more complex S-shaped or double-curvature patterns.

7. More flexibility creates more identifiability risk

If τ1 and τ2 become similar, L2 and L3 become nearly collinear.

Then very different β2,β3 combinations can create nearly the same curve.

The fitted yields may be stable while the parameters are unstable.

Falsifier: monitor the singular values or condition number of the loading matrix. Parameter instability with tiny curve changes is evidence of weak identification, not necessarily a bad numerical optimizer.

8. Conditional on decay parameters, the betas are linear

For fixed τ1,τ2, the loadings Lj(τi) are known numbers.

The β parameters can then be estimated by linear least squares:

β̂=(XTWX)−1XTWy

when fitting zero yields with weights W.

This means NSS is a separable nonlinear least-squares problem: only the decay parameters need nonlinear search; the linear betas can be solved conditionally.

This structure can make calibration faster and easier to diagnose than treating all six parameters as one black-box nonlinear optimization.

9. Grid search or multi-start for decay parameters

Because the objective can contain multiple local minima and flat regions, a single optimizer starting point is risky.

Useful approaches include:

  • bounded grid search over τ1,τ2 followed by local refinement;
  • multi-start nonlinear optimization;
  • profile objective plots;
  • variable-projection/separable least-squares algorithms;
  • constraints keeping τ parameters apart when justified.

Falsifier: run the calibration from many starting points. If equally good fits return radically different parameters, report the identification problem instead of publishing one parameter vector as unique truth.

10. Yield-error versus price-error calibration

Suppose the inputs are coupon bonds.

One approach converts each bond to a yield and minimizes:

Σ wi[ymodel,i−ymarket,i]².

Another directly prices each bond from the model discount curve and minimizes:

Σ wi[Pmodel,i−Pmarket,i]².

These objectives are not equivalent because a one-basis-point yield error creates different price errors for short and long-duration bonds.

The calibration objective should match the intended use and weighting philosophy.

11. Duration weighting

If price errors are minimized without weighting, long-duration bonds can dominate because their prices are more rate-sensitive.

If yield errors are minimized, short/long securities may receive similar basis-point importance.

Possible weights include:

  • equal yield-error weights;
  • inverse bid–ask variance;
  • duration/DV01 normalization;
  • liquidity weights;
  • issue-size or maturity-bucket weights.

Weights encode a loss function. They are not a harmless implementation detail.

12. Coupon bonds require cash-flow pricing

A coupon bond cannot be matched by reading one curve point at its maturity.

Its full dirty price is:

P=Σ CjD(tj) + FD(T),

where D(t) is the model discount factor.

With continuously compounded zero yield y(t):

D(t)=e−y(t)t.

The calibration must use correct coupon dates, day counts, accrued interest and settlement conventions.

Falsifier: reprice every calibration bond from its full cash-flow schedule. A curve that fits quoted yields but cannot reproduce prices under the actual conventions is not correctly implemented.

13. Clean versus dirty prices

Market coupon bonds may be quoted clean but settled dirty:

Pdirty=Pclean+accrued interest.

Calibration code must compare like with like.

A systematic maturity-pattern residual can be nothing more than an accrued-interest/day-count error.

14. Instrument selection matters

Official curve methodologies often exclude securities whose prices contain strong non-term-structure effects.

For example, the Federal Reserve’s current nominal Treasury curve excludes on-the-run and first-off-the-run notes and bonds because those issues can carry liquidity/repo specialness premiums. It fits a smoother off-the-run term structure.

The ECB likewise applies security-selection and outlier rules to government bonds before fitting its reference curves.

Public lesson: a curve model cannot distinguish “interest-rate shape” from “security-specific cheapness/richness” unless the data pipeline does.

15. Institutional use is evidence of practicality, not universal optimality

The Federal Reserve currently publishes daily nominal yield-curve parameters using a Svensson specification for the modern Treasury sample and Nelson–Siegel for earlier periods with fewer securities.

The ECB publishes Svensson-model parameters for euro-area government-bond curves under documented conventions.

These are strong demonstrations that NSS is operationally useful.

They do not mean every asset class, credit curve or derivatives discount curve should use NSS rather than bootstrapping/interpolation.

16. Spot, forward and discount curves

From y(τ), continuously compounded discount factors are:

D(τ)=e−τy(τ).

The instantaneous forward rate is:

f(τ)=d[τy(τ)]/dτ.

Under the NSS form:

f(τ)=β0+β1e−τ/τ1+β2(τ/τ1)e−τ/τ1+β3(τ/τ2)e−τ/τ2.

Forward curves often expose fitting problems more clearly than spot curves because derivatives amplify local wiggles.

17. Residual diagnostics

Do not report only overall RMSE.

Plot residuals by:

  • maturity;
  • security;
  • coupon;
  • liquidity category;
  • issue age;
  • bid–ask spread;
  • time.

Repeated positive residuals around one maturity can indicate a missing curve shape, a market segmentation effect or data/convention errors.

18. Forward-rate diagnostics

A visually smooth spot curve can hide implausible forward-rate behavior.

Check:

  • large forward spikes;
  • rapid oscillations;
  • negative forwards when economically/methodologically unexpected;
  • long-end convergence;
  • day-to-day forward instability.

Negative rates are not automatically an arbitrage violation; they have existed in real markets. The diagnostic is consistency with discount factors, instruments and economic context—not a hard “rates must be positive” rule.

19. Discount-factor monotonicity

For a standard positive-discounting environment, discount factors commonly decrease with maturity. But negative interest rates can make D(T) exceed one and even rise over some ranges without creating a simple static arbitrage under the relevant numeraire/conventions.

Therefore naïve positivity/monotonicity rules must reflect the market regime.

A better no-arbitrage audit begins with whether model discount factors and forwards generate internally consistent prices for admissible cash flows under the chosen framework.

20. Static smoothness is not dynamic no-arbitrage

An NSS curve can fit today’s cross-section perfectly well.

If its parameters are then given arbitrary time-series dynamics, the resulting stochastic term-structure model need not be arbitrage free.

Christensen, Diebold and Rudebusch explicitly developed an arbitrage-free generalized Nelson–Siegel framework because ordinary dynamic Nelson–Siegel/Svensson factor evolution does not automatically impose finance no-arbitrage restrictions.

This page owns cross-sectional fitting; it does not turn the six parameters into a risk-neutral derivative-pricing model.

21. NSS versus bootstrapping

Yield-curve bootstrapping algorithms recursively solve discount factors so selected market instruments reprice exactly, then interpolate between knots.

NSS instead accepts cross-sectional residual error in exchange for a compact smooth global curve.

For collateralized derivatives discounting, exact instrument reproduction with local interpolation is often essential. For macro/reference-curve publication, a smooth NSS fit can be preferable.

22. NSS versus Hagan–West monotone-convex interpolation

Hagan–West monotone-convex algorithms own local interpolation between bootstrapped forward-rate knots while trying to preserve desirable forward behavior.

NSS is a global parametric fit. A quote at one maturity can influence the fitted curve far away through the shared parameters.

23. NSS versus historical PCA

Both methods often use the words level, slope and curvature, but they are mathematically different.

NSS loadings are pre-specified exponential functions and parameters are fitted to one cross-section.

PCA loadings are data-driven eigenvectors estimated from a time series of curve changes.

Similar names do not imply identical factors.

24. The role of τ parameters

τ1 and τ2 control the decay scales of slope/curvature loadings.

Very small τ can make factors concentrate at the extreme short end.

Very large τ can make loadings nearly collinear over the observed maturity range.

Bounds should therefore reflect the maturity span and data resolution rather than arbitrary optimizer defaults.

25. Short-end identification

If the shortest observed instrument is one year but τ1 is intended to control behavior inside the first few months, the data contain little direct information about that feature.

β1 and τ1 can trade off to produce similar one-year-plus curves.

Falsifier: inspect profile objectives and parameter confidence regions. Do not claim a precise short-end factor when no short-end instruments exist.

26. Long-end identification

β0 is the infinite-maturity asymptote, but no market observes infinity.

If the longest security is 10 years, many combinations can extrapolate differently beyond 10 years while fitting the observed sample similarly.

Falsifier: extend the curve to 20, 30 and 50 years under alternative near-optimal fits. Wide divergence means long-end extrapolation is weakly identified.

27. Second-hump overfitting

Svensson’s β3,τ2 pair can fit a genuine second curvature feature—or chase noise.

If Nelson–Siegel already fits within market bid–ask noise, adding the second hump may create unstable parameters with little economic benefit.

Falsifier: compare out-of-sample price errors and forward-curve stability, not only in-sample RMSE.

28. Daily parameter stability

A published curve can be numerically smooth each day while parameters jump wildly from one day to the next.

This often indicates:

  • weak identification;
  • multiple local minima;
  • changing instrument sample;
  • outliers;
  • τ-factor collinearity.

Parameter jumps matter when downstream systems interpret them as economic factor changes.

29. Temporal regularisation

One can add a penalty encouraging today’s parameters to remain near yesterday’s:

Loss = fit error + λ||θt−θt−1||².

This can improve stability but introduces path dependence and deliberate bias.

A reference curve intended to represent today’s market independently may prefer unregularized fitting plus diagnostics instead.

30. Outliers and liquidity

A single stale/illiquid bond can pull a global parametric curve across many maturities.

Controls can include:

  • bid–ask screens;
  • issue-quality rules;
  • robust residual checks;
  • iterative outlier review;
  • minimum securities per maturity region;
  • comparison with local/nonparametric curves.

The ECB’s public methodology, for example, describes security selection and repeated outlier removal for its government-bond curves.

31. Inputs and outputs

Inputs can include:

  • bond/zero-rate observations;
  • settlement and cash-flow conventions;
  • clean/dirty prices;
  • maturity/coupon/liquidity data;
  • weighting rule;
  • Nelson–Siegel or Svensson model choice;
  • τ parameter bounds;
  • optimizer/multi-start settings;
  • outlier policy;
  • compounding convention.

Outputs can include:

  • β0…β3;
  • τ1,τ2;
  • spot curve;
  • discount factors;
  • instantaneous forward curve;
  • par rates;
  • instrument repricing errors;
  • weighted RMSE;
  • condition/identifiability diagnostics;
  • alternative near-optimal fits.

32. Evidence polarity

Evidence for confidence includes:

  • instrument residuals lie within reasonable market noise;
  • multiple optimizer starts converge to similar curve shapes;
  • loading matrix is well-conditioned;
  • forward curve is smooth and economically interpretable;
  • out-of-sample bond repricing is stable;
  • parameters are reasonably stable or parameter instability is shown not to affect the curve;
  • long-end/short-end extrapolation is supported by actual observations;
  • results agree with an independent curve method over liquid maturities.

Evidence against confidence includes:

  • many near-equal minima with different curves;
  • τ1≈τ2 and severe collinearity;
  • large maturity-pattern residuals;
  • forward-rate spikes;
  • parameters jump while market curves barely move;
  • second curvature term improves only in-sample noise fit;
  • long-end extrapolation changes dramatically under near-optimal calibrations;
  • price conventions or security-specific premiums contaminate the fit.

33. Counterexample: excellent RMSE, unstable parameters

Two parameter sets produce yield RMSE below one basis point but β2,β3,τ1,τ2 differ dramatically.

Falsifier: compare the fitted curves and the loading-matrix singular values. The curve may be fit well while the decomposition is not identified.

34. Counterexample: fitting an on-the-run liquidity premium

A newly issued Treasury trades rich because of liquidity/repo specialness.

A global NSS curve bends toward it and makes surrounding off-the-run bonds appear cheap.

Falsifier: refit after excluding known special/liquid issues. A large curve shift indicates the original fit mixed term-structure and liquidity effects.

35. Counterexample: second hump chases one bad bond

One illiquid 12-year bond is mispriced. Svensson uses its extra curvature factor to fit that point, creating a forward-rate bump around the same region.

Falsifier: leave the bond out and refit. If β3,τ2 collapse and nearby liquid instruments price better, the second hump was fitting noise.

36. Counterexample: extrapolation treated as observation

The longest bond is 15 years, but a 50-year NSS yield is reported with the same confidence as a five-year yield.

Falsifier: produce extrapolation bands from near-optimal parameter sets or alternative curve families. Confidence should decay outside the supported maturity region.

37. Counterexample: static fit used for derivative dynamics

Daily NSS parameters are each modeled as independent Gaussian AR processes and used to price long-dated options.

Good cross-sectional curve fit does not guarantee that these parameter dynamics are arbitrage consistent.

Falsifier: test whether the resulting discount-bond dynamics admit an equivalent martingale measure/no-arbitrage structure. If not, use a dedicated arbitrage-free term-structure model.

38. Alternatives

Bootstrapped piecewise curves: exact reproduction of calibration instruments with local interpolation.

Monotone-convex interpolation: controlled forward-rate shape between knots.

Splines: flexible local/global smoothers with different oscillation risks.

Smith–Wilson: regulatory long-end extrapolation toward an ultimate forward rate.

Arbitrage-free affine term-structure models: dynamic pricing models rather than purely static smoothers.

39. Weak links

  • wrong formula convention for τ versus λ;
  • incorrect short-maturity limit handling;
  • cash-flow/day-count/accrued-interest errors;
  • τ bounds inconsistent with maturity range;
  • single-start optimizer;
  • τ-factor collinearity;
  • liquidity/outlier contamination;
  • price/yield weighting mismatch;
  • over-interpreting unstable β parameters;
  • treating static fit as dynamic no-arbitrage model.

40. What would falsify confidence?

Confidence should be withdrawn if calibration cannot reprice clean benchmark securities within reasonable error; if multi-start solutions disagree materially; if the loading matrix is nearly singular; if forward curves develop unexplained oscillations; if extrapolated maturities are highly unstable; or if the model is used dynamically without separate no-arbitrage validation.

41. Verification and update triggers

Preserve the exact formula convention, compounding basis, instrument universe, cash-flow engine, weights, τ bounds, optimizer starts, residuals, condition diagnostics and alternative fits.

Revalidate when:

  • instrument universe changes;
  • market liquidity changes materially;
  • negative-rate regime/conventions change;
  • τ parameters hit bounds;
  • condition numbers deteriorate;
  • forward-rate residuals/spikes appear;
  • new long-end securities extend the observable curve;
  • the curve is repurposed for a new pricing/risk job.

42. Primary and high-quality references

Educational boundary: Nelson–Siegel–Svensson is a compact curve-fitting language. A beautiful six-parameter fit is not proof that the parameters are uniquely identified, that extrapolated maturities are observed, or that arbitrary parameter dynamics form an arbitrage-free pricing model.

Discover more from Bukit Timah Tutor

Subscribe now to keep reading and get access to the full archive.

Continue reading