Skip to main content

Teaching Valhalla That Hiking Is Not Flat

· 9 min read
Evgen Bodunov
Evgen Bodunov
CEO @ Globus Software

A walking route across a city and a hiking route up a mountain should not share the same clock.

Routing engines are good at distance, turn costs, access rules, and road classes. Hiking time is harder. A two-kilometer trail can be easy, slow, dangerous, or all three, depending on slope, descent, surface, visibility, and how the local trail network publishes reference times.

Globus routing is built on Valhalla, and Valhalla tiles already contain the graph, the route shape, and elevation. We wanted to know whether that information was enough to make hiking ETAs less optimistic, without breaking tile compatibility and without a formula that only works in one mountain range.

We built a corpus of reference hiking times, tested several formulas, rejected most of them, used a deliberately expensive research hook to find what worked, and then moved the result into production Valhalla code.

The problem​

Flat walking speed is not a hiking model.

On a paved path through a park, distance explains most of the time. On a route that climbs 700 meters, distance alone predicts badly. On a steep alpine descent, even "downhill is faster" stops being true.

The usual failure is optimism: the route looks short on the map, so the ETA is short in the app. Anyone who has followed a marked mountain route knows that the posted time is often more honest.

We did not set out to build a perfect hiking guide. The goal was narrower:

  • use information Valhalla already has
  • preserve old-tile and new-code compatibility
  • avoid a formula that only works in Poland, Switzerland, or Italy
  • keep the model explainable enough to debug
  • measure against real reference hiking times, not intuition

The data we used​

We started in Lesser Poland, where many OSM hiking relations include reference durations in both directions. These are useful because the same trail can have different uphill and downhill times.

Then we added Switzerland. Swiss hiking times come from a public signpost system with a long mathematical tradition. The exact coefficient table is not public, but the official guidance is clear: time depends on horizontal distance and slope, with special care for very steep sections.

Finally, we added an Italian Alps stress corpus with more short, steep, difficult routes. It also served as a separate check on whether a model tuned on Poland and Switzerland had learned hiking time or just local habits.

The final comparison used exact relation-shape elevation profiles where possible. Routing between endpoints is useful for service testing, but for model validation it can add detours that are not part of the published hiking relation. Validating on the exact shape is cleaner: the prediction is made for the same geometry that owns the reference duration.

What did not survive​

The first idea was accumulated elevation gain and loss.

It gets you much further than distance alone. It knows that a route with 600 meters of ascent is not a normal walk, and it preserves direction: the uphill and downhill versions of the same trail can differ.

But gain and loss alone discard too much. They do not distinguish a steady ascent from steep steps, or mild downhill, which can be fast, from steep downhill, which can be slow. They also react badly when noisy elevation samples inflate tiny ups and downs.

We also tested difficulty tags such as sac_scale. The signal is real on some Polish high-alpine segments, but it overfits easily. Some Swiss routes with high difficulty tags were already predicted too slow by slope-based models, and a generic technical penalty made them worse.

We tested Tobler-style walking functions. Their shape is useful: mild downhill is fastest, climbing is slower, and steep descent eventually becomes slow too. The literal formula, however, was too optimistic for our combined data. Tuned Tobler variants did better, especially in Switzerland, but were not the best single answer across all corpora.

We tested Weber-like slope polynomials inspired by the Swiss signpost method. They performed very well, and one blended polynomial had the best aggregate MAE in our clean check set. The problem was production: high-order polynomial coefficients are hard to reason about and easy to destabilize outside the fitted range.

What we took away was not a particular formula but a shape: hiking ETA needs a slope-speed curve.

The model we kept​

We kept a bounded slope curve.

During research it ran at path level. Valhalla collected the elevation profile along the routed path, smoothed slope over a 150-meter window, and assigned seconds per kilometer from a small piecewise curve:

slope band: -40% -20% -8% -5% +3% +8% +20% +40%
seconds/km: 2520 1270 1130 880 1020 1110 2330 4120

The shape is deliberately simple:

  • mild downhill is fastest
  • steep downhill slows down
  • climbing slows down as grade increases
  • extreme slopes are clamped instead of extrapolated
  • the same formula runs in every region

It is less elegant than a fitted polynomial, but easier to inspect, to bound, and to explain when a route is wrong.

The path-level version suited research because it saw the whole route. It did not suit production: Valhalla routes are made of directed edges, and costing has to be fast. Once the model looked good, we checked whether the same curve still behaved well when computed edge by edge, using 150-meter chunks inside each graph edge and a shorter final chunk when needed.

The numbers​

After quality filtering, the clean check set had 250 directional route samples across Poland, Switzerland, and Italy.

The best aggregate model was the Weber/global blend:

clean combined check MAE: 7.532 minutes
clean combined check MAPE: 13.312%
bias: -1.694 minutes

The selected bounded slope curve was almost tied:

clean combined check MAE: 7.578 minutes
clean combined check MAPE: 13.135%
bias: -2.123 minutes

We gave up 0.046 minutes of average absolute error, roughly three seconds per route, for a model with a clearer physical shape and slightly better percentage error. Routing models outlive benchmark scripts, and a tiny metric win does not justify a formula that is harder to debug, harder to port, or easier to break with new data.

The remaining errors were informative as well. Quality-flagged rows had roughly 30 minutes MAE whichever model we used. Some were probably source-format issues: one Swiss relation had a bare numeric duration 334, which as minutes means 5:34 but may have been intended as compact 3:34. Others had elevation mismatch between the route profile and the source ascent and descent.

In other words, some outliers are not model problems. They are corpus problems, geometry problems, or missing machine-readable trail difficulty.

The production check used exact edge traces from Poland and Switzerland. On the clean check group, 150-meter segments computed edge by edge produced:

clean check MAE: 8.361 minutes
clean check MAPE: 11.053%
bias: -2.590 minutes

The path-level and edge-local versions were close enough to proceed with the faster edge-local implementation. On the combined check set used for this comparison the edge-local version was slightly better, and, more importantly, it can be precomputed in tiles.

Compatibility​

We wanted compatibility from the start.

The research hook did not require new tiles. It decoded the elevation already present in existing Valhalla tiles and computed the path-level slope curve while building the route response, so we could iterate without rebuilding the planet.

The production implementation keeps that property but moves the work to tile generation. Valhalla already has an extended directed-edge block that can carry extra per-edge attributes without changing the base DirectedEdge layout. We use it to store one precomputed hiking duration per directed edge.

The storage is small:

  • 16 bits for hiking seconds
  • one validity bit
  • remaining bits reserved for future extension

Two bytes hold up to 65,535 seconds per directed edge, more than 18 hours, which is enough for real graph edges. If tile generation ever computes more, it logs a warning and clamps the stored value so the edge can be investigated.

As a result:

  • old code can still read new tiles because the base edge layout is unchanged
  • new code can still read old tiles because missing extension data falls back to the old pedestrian timing formula
  • the new model is opt-in through pedestrian type=hiking
  • normal foot, wheelchair, and blind pedestrian behavior stays unchanged

Edges are directed, and that matters here: the two directions of the same trail can have different hiking times, because uphill and downhill are not symmetric. Tile generation may compute forward and reverse values while processing a shared shape, but each directed edge stores only the value for its own orientation.

What this means for routing​

The result is less a coefficient table than a set of findings about hiking ETA:

  • distance-only walking is too optimistic for mountain routes
  • raw accumulated gain/loss helps but is not enough
  • generic difficulty penalties overfit quickly
  • slope-window models transfer better across countries
  • quality flags are necessary before trusting outliers
  • the production formula should be bounded and inspectable

For users, flat walks should stay close to normal walking time, moderate downhill can be faster, and steep climbs and descents no longer look like short city walks.

For Valhalla deployments, hiking ETA can improve using existing elevation data, without a breaking tile-format migration. The production version computes hiking seconds during elevation building and reads them during pedestrian costing when the request uses type=hiking.

What comes next​

The research is finished and the production code is in place.

Hiking ETA will never be as clean as road speed. Trails carry local timing conventions, ambiguous geometry, incomplete difficulty tags, and human caution. But it can be much better than flat walking speed, with a change small enough to explain, measured against real routes, and compatible with the map data people already have.

To try the hiking profile, create an API key. For questions about Valhalla navigation tiles or custom routing deployments, write to [email protected].