Foundation matching fails more often than almost any other purchase in beauty. The reason is that skin varies along several independent dimensions while most shade ranges are built along one.
Depth and undertone are separate measurements
Depth is how light or dark skin is. Undertone is the colour cast underneath, usually described as warm, cool or neutral, and it is independent of depth.
A range organised primarily by depth forces undertone into a fixed pattern, typically assuming warmer tones as shades get darker.
Anyone whose undertone does not follow that assumption finds shades that are the right lightness and the wrong colour, or the reverse, with nothing between them.
Skin is not a single colour across the face
The centre of the face is usually redder from underlying vasculature, the jaw and neck are often lighter, and the forehead may be deeper from sun exposure.
A single foundation shade averages across all of that, so it will sit correctly somewhere on the face and slightly wrong elsewhere.
Matching along the jaw rather than on the back of the hand is standard practice because the jaw sits between face and neck, where a visible line would otherwise form.
Pigments interact with light differently from skin
Foundation colour comes mostly from iron oxides and titanium dioxide, which scatter and absorb light in their own way rather than reproducing skin's translucency.
Skin transmits some light beneath the surface before it returns, giving a glow that an opaque layer of pigment sitting on top cannot reproduce.
This is why a shade that matches perfectly in the bottle can read flat or mask-like on the face even when the colour measurement is correct.
Formulas oxidise after application
Many foundations shift colour in the first hour on skin, usually toward a deeper or more orange tone. The effect is commonly called oxidation.
It results from the formula interacting with skin oils, with air, and with whatever sits underneath it, so the same product shifts differently on different people.
A shade chosen at the moment of application can therefore be the wrong one an hour later, which is why testing is best judged after a wait rather than immediately.
Range breadth is a manufacturing decision
Every additional shade is a separate batch, a separate stock line and a separate risk of unsold inventory, and demand across a range is deeply uneven.
Historically this pushed brands toward narrow ranges clustered around the shades that sold fastest, leaving the extremes of depth poorly served.
Ranges have widened substantially in recent years, and the improvement is mostly in undertone variety at each depth rather than in the number of shades alone.