BETA — measurements and names under active refinement. Seen a sign, a gate, or a better name? Submit an update no account needed · photos welcome via GitHub

← survey home · UK climbs · US climbs

How to read these numbers

Why every maximum states a distance

The gradient of a road depends on the distance you measure it over. The same hill can honestly be 32% over 2 m, 29% over 5 m and 25% over 100 m — all three at once. A “maximum gradient” without a distance is unfalsifiable, which is why signs, cycling databases and GPS platforms so often disagree: they are answering different questions. Every value here is a window maximum: the steepest stretch of exactly that plan-distance anywhere on the climb.

Why 2 m and 3 m are “diagnostic”

A window gradient is computed from two terrain samples, so elevation noise propagates as σ ≈ √2 × σelev / window. The source surveys quote ±15 cm RMSE absolute vertical accuracy; what a short window differences is the relative error between two nearby cells, which is smaller because systematic components (datum, sensor calibration, strip adjustment) are largely shared between neighbours — we use ±10 cm per sample as a working estimate of local relative error, one the planned calibration exercise is designed to test. With that, the raster-noise uncertainty alone is:

Window2 m3 m 5 m10 m25 m 50 m100 m
1σ noise±7.1 pp ±4.7 pp±2.8 pp ±1.4 pp±0.6 pp ±0.3 pp±0.14 pp

Two caveats on that table. It assumes the endpoint errors are independent; bilinear interpolation and a shared LiDAR point cloud correlate nearby samples, which damps random noise but means correlated errors (canopy, road benches, embankments) are not covered — so read it as a floor for random noise, not a total error budget. And at 2 m the noise is as large as the differences being claimed, with the systematic errors worst there too: a metre of centreline placement error moves the line onto road camber or verge, and interpolation smooths features shorter than about two raster cells. There is also a physical point — 2 m is only around two bicycle wheelbases (a wheelbase is roughly 1 m), so a 2 m maximum is a ramp the bike momentarily spans, not a sustained pitch anyone rides.

We publish 2 m and 3 m values anyway, greyed and labelled diagnostic, because that is where headline claims live: gradient signs, platform “max grade” figures and record adjudications often appear to reflect short sections on roughly the 2–10 m scale. Diagnostic windows exist to interpret those claims. They are never used in rankings. For the same reason, windows of 10 m and below are displayed as whole percentages — at ±1.4 pp or more of noise, a decimal place would be faux precision. Stored OCL documents keep full precision.

Quality interventions

Three tiers, in increasing force — raw data is never modified:

≈ the surface model sits well above the terrain along this window (tree canopy, walls, buildings), so the bare-ground model there is built from fewer laser returns: the value stands but is less certain — rankings and ladders mark it with ≈.
excluded the window crossed ground where the terrain model does not represent the road: a sudden break steeper than any road surface, a mapped bridge or tunnel, the terrain shadow under a road crossing overhead, a steep V-shaped dip with no road-like floor, dense plantation whose cross-sections show no road bench at all (the surveyed “ground” there is canopy), ground so steeply banked across the line that no carriageway surface survives in the model at all — a sunken, walled or sharply embanked lane, where what the model holds is bank, and the gradient read there is a function of where the line was drawn rather than of the road — or a span a human reviewer has adjudicated (recorded in the climb’s provenance). The maximum shown is the steepest clean stretch and the displaced raw reading is noted.
needs review the final numbers still trip plausibility checks and await human scrutiny.

How the numbers are policed

A “steepest” ranking has an adversarial property: the steepest thing in any dataset is disproportionately likely to be an error, because errors are unbounded and roads are not. We treat our own headline values accordingly. The masks above are the first line of defence; the second is a battery of detectors that sweeps the whole corpus after every batch of new or re-measured climbs. Each detector encodes a failure we actually found — usually at the top of a table:

The detectors only report — none of them changes a number. A person adjudicates every flag; verdicts are recorded and settled, and the detectors consult those verdicts so a human-confirmed measurement is never re-litigated by automation. Whoever tops a ranking table after any change to the dataset gets a fresh human check — three different crowns fell to it in a single night. Extent gets audited too: every climb’s endpoints are re-walked to ask whether the road keeps climbing beyond them. And retirements are public: the climb count goes down when a measurement fails scrutiny (one 2026 measurement-integrity audit retired some 95 pages). An error we published and then caught is removed, not quietly patched.

Where a climb is allowed to end

Candidate routes are grown along the road network, and the growth sometimes runs a step too far — onto a private drive, a gated track or an unnamed service way at the base or the summit. For a long time one such appended segment condemned the whole candidate: the access audit judges the full route, so the climb was rejected outright and never measured. That was the right verdict about the segment and the wrong granularity about the road. The survey now cuts the route at the access boundary and lets the public remainder stand or fall on its own numbers; remainders too short to qualify are dropped, and a page produced this way records the cut — where and how much — in its provenance. The steep private section it lost is not pretended away: it simply is not the public climb. Access questions that a map cannot settle (a gate that is really a farm courtesy, a “private” lane a council resurfaces) are judged by a person, and the verdict is recorded and settled.

How the US survey differs

The US section applies the same windows, masks and confidence rules to different national data, and two things about rural America change the audit posture.

Elevation. US climbs are measured on USGS 3DEP 1 m LiDAR bare-earth models (NAVD88) — the same resolution bar as the UK survey. A coarser 10 m model is used only to find candidate climbs; published numbers always come from the 1 m data.

Canopy. 3DEP publishes no national surface model, so the canopy check is measured from the survey’s own point cloud along each route corridor: the highest laser return over each metre of road, against the bare-earth model from the same flight. Where no point cloud is published yet (parts of Vermont and western Maine await release), the page says the canopy check is unavailable — absence of evidence is recorded, never converted into “clear”. Most US LiDAR is flown leaf-off, which is also why a missing canopy check matters less there than it would in the UK.

Roads. There is no US equivalent of the OS road-network centreline, and rural OpenStreetMap inherits old imports — including mapped ways with no road under them. So US rows are verified against the state road inventories (VTrans, NH GRANIT, MaineDOT, MassDOT, CTDOT, RIGIS): every ranked row carries a surface chip (paved / mixed / gravel, with a paved-only filter), and the gating window itself — the 25 m stretch that earns the rank — is checked against the inventory, not just the route as a whole. A row whose gating window the inventory calls unmatched, unimproved or private moves to a “road status under review” table instead of the rankings; private and toll climbs rank alongside everything else with a controlled-access note on the row and the page — a real road lists, with its restriction stated, never silently ranked away. The same policy holds survey-wide: an access-restricted or rights-uncertain climb appears in the lists with its caveat.

Existence. Because a mapped way is not proof of a road, an audit asks the point cloud whether a graded road bed physically exists along the line: a real road is a bench — a smooth strip that stays cross-level while the hillside it is cut into tilts — and untouched terrain does not do that. The check is calibrated against human-verified cases before it is allowed to flag anything, it deliberately returns “inconclusive” on flat ground (where geometry alone cannot distinguish a road bed from smooth forest floor), and like every detector it only reports: a human adjudicates every flag before a page changes.

Measurement confidence — the A/B/C/D/U letters

Every climb carries a single letter summarising how much trust to place in its numbers. It is deliberately not called a “grade” or a “category” — in cycling both of those words mean the gradient itself. The letter says nothing about how steep the hill is; it describes the evidence:

High (A) — 1 m LiDAR, standard windows clean: no canopy flags, no re-sited maxima, nothing unresolved.
Good (B) — sound but with minor caveats: a source finer than 2 m but coarser than 1 m, or 1 m data where standard windows needed intervention (canopy ≈ flags, re-sited maxima).
Limited (C) — a source coarser than 2 m up to 5 m, or standard windows that could not be cleanly placed after artifact exclusion.
Low (D) — a source coarser than 5 m; short-window maxima suppressed.
Needs confirmation (U) — the numbers trip our own plausibility checks and await human review — or your local knowledge: treat as provisional.

The letters are what the machine-readable OCL documents carry (compact and stable); these pages show the words. Only High, Good and Needs confirmation occur in the current dataset — Limited and Low exist for future regions with coarser elevation data.

Rankings on the front page include confidence A/B only; U climbs are listed separately as provisional discoveries. Two things deliberately do not affect the letter: extent questions (“should the climb start earlier?” — they concern the climb’s definition, not the measurement), and ≈ flags on the 2–3 m diagnostic rows — those windows are never headline values, so a climb can honestly show ≈ there and still carry high confidence in the numbers that matter.

What “Matches” means

The Matches column records external evidence matched to OUR measured object — lists that include it, and documentation physically on the road. Matching is best-effort and ongoing; the measurement never depends on it:

Listed — the climb appears on an external list we record (books, championship venues, race routes), cited by numbered footnote. Membership facts only: we never copy a list’s own measurements.
sign: 32% (2025) — a physical gradient sign documented on the road, with the year of the observation. Signs get repainted (one of our climbs went 25% → 32% between observations), so the date is part of the fact. Where the sign is known through OpenStreetMap’s record of it rather than a direct observation, the entry says “via OSM” — strong evidence, but not a photographed sign.
tag: 25% (2026) — an OpenStreetMap incline value or mapper estimate, dated by the map snapshot we read it from. Weaker evidence than a sign: usually right, but some tags are impressions rather than sign readings — our measurements test them either way.
— — not yet matched against documented climbs. The survey measures every road meeting its criteria from mapping and LiDAR data alone; matching those measurements against books, race records and lists is a separate, ongoing, best-effort layer — an unmatched climb may be famous locally. The matching is also something anyone can do with our open data, and reports of known names are the quickest way a climb gets matched.

Seen a sign we don’t show, or one that has changed? Submit an update (no account needed; photos welcome via the GitHub option there) — on-the-ground observations are evidence we record and test, exactly like the signs and map tags already in the dataset.

What these numbers are not

They are terrain-model measurements along the mapped road centreline — not asphalt surveys. All distances are plan (map) distances measured along the centreline's path: on a hairpin, a 5 m window follows the curve of the carriageway, never a straight line across it, and route lines are audited against a national road-network centreline (Ordnance Survey in the UK, the state road inventories in the US) so a window can't quietly measure the hillside inside a bend.

The measurement’s authority is the data and the method: a transparent algorithm applied to professionally flown, quality-controlled national elevation surveys, with the uncertainties stated above. Independent observations — signs, ride reports, spot checks — are recorded as evidence and tested against the measurement (that is the Matches column above); where they agree, the corroboration is noted on the page. A point observation cannot re-derive a windowed average, so evidence annotates a measurement rather than outranking it.

Separately, an uncertainty-calibration exercise is planned: profiling a small set of deliberately varied test roads — clean terrain, tree canopy, walls and buildings, extreme gradient — with independent field measurement, to test the ±10 cm working estimate the uncertainty table above rests on. Calibration tightens the stated error bars for every climb at once; it is a check on the method’s error model, not a per-climb blessing. Short-window values are, and will remain, estimates with stated uncertainty — that is what honest measurement at this scale looks like.