AI Running

What the “AI” in Your Running App Is Actually Doing

Most of the apps selling you an AI coach are running a spreadsheet with a nice onboarding flow. That is the short version. The longer version is more useful, because a small number of them genuinely are modelling you, and the difference shows up in your legs about seven weeks into a marathon block.

Here is the distinction that matters. A parameterised template takes a handful of inputs at signup (goal race, date, days per week, current 5K time, maybe a “how hard do you want to go” slider), picks a plan shape from a library, and fills in your paces with arithmetic. It’s a very good plan generator. It is not watching you. An adaptive model carries a running estimate of your state, updates that estimate every time a file syncs from your watch, and recomputes what you should do next from the updated numbers. The first one answers “what does week 9 look like for a 4:15 marathoner.” The second answers “what should this person do on Tuesday given what they actually did for the last 28 days.”

Both can produce a sensible Tuesday. Only one of them notices when you stop being the person who signed up.

How do AI running apps work, mechanically

Strip the branding and there are three layers, and you can usually work out which ones an app has.

Layer one is pace and zone derivation. You give it a recent race or time trial, it derives training paces. Riegel’s formula for race equivalents, something like Daniels’ VDOT or Jack Tupper’s tables for zone boundaries. A 22:30 5K gives a threshold pace near 4:35/km, an easy pace near 5:30 to 5:50/km, a marathon pace near 5:00/km. This is a lookup. Every app has it. It is not AI in any sense a statistician would recognise, and it was fully solved by 1978.

Layer two is load accounting. The app converts each session into a single fatigue number and tracks two moving averages of it: something acute (7 days) and something chronic (28 to 42 days). TrainingPeaks calls these ATL and CTL and the ratio TSB. Garmin calls them acute load and chronic load and exposes the ratio directly on newer Forerunner and Fenix watches. Coros calls the chronic number Base Fitness. Intervals.icu, which is free, calls them Fatigue, Fitness and Form. These are the same family of model, all descended from Banister’s impulse-response work from the 1970s, and they are genuinely predictive at the crude level of “your chronic load went up 40% in three weeks, which historically precedes people getting hurt.”

Layer three is closing the loop. The app changes what it prescribes based on layer two, without you asking. This is the rare one.

Plenty of apps have one and two and stop. They’ll show you a lovely fitness curve, and then serve you the Tuesday session that was written in the plan on day one regardless of where that curve went. A dashboard is not a controller.

The test that separates them

Miss a week. Not deliberately: it will happen anyway, so watch what the app does with it.

Say you’re in week 9 of a 16-week marathon build and you get a chest infection on the Sunday of week 8. Six days off, zero running. Your plan as originally written for week 9:

Tue   6 x 1000m @ 4:05/km, 90s jog recovery
Thu   8 km easy @ 5:30/km
Sat   24 km long @ 5:20/km
Sun   10 km easy
Week total: 48 km

Meanwhile, in the load layer, this happened:

Fitness (CTL)    42  ->  31     (-26% over 12 days)
Fatigue (ATL)    38  ->   9
Form (TSB)       +4  ->  +22
7d load / 28d load ratio    1.05  ->  0.31

A template app serves you that 48 km week unchanged, because the week is a row in a table and the table does not read your files. If you run it, your 7-day load goes from near zero to a chronic-load ratio around 1.6, which is well outside the 0.8 to 1.3 band most of the injury literature treats as boring and safe. You are also being asked to hold 4:05/km off six days of nothing.

A model-driven app does something closer to this: it drops the week to roughly 36 to 38 km, replaces the 6 × 1000m with 4 × 1000m at 4:12 or a fartlek with looser targets, shortens the long run to 18 km, and pushes the original week 9 content to week 10. It does that because it recomputed from the new CTL, not because you told it you were ill.

Garmin is an interesting case here, because it ships both kinds under one roof. Garmin Coach is a template: real plans, written by real coaches (Greg McMillan, Amy Parkerson-Mitchell, Jeff Galloway), with limited responsiveness to what you actually did. Daily Suggested Workouts on the same watch are state-driven: they read acute load, chronic load, HRV status and recovery time, and they visibly change their minds. Come back from that infection and the suggestion is a 40-minute base run, not intervals. Same manufacturer, same watch, two completely different philosophies, and nothing in the marketing tells you which one you picked.

Stryd is worth naming for the opposite reason: it’s honest about the mechanism. Its plans compute from your Critical Power, which updates from your runs. CP moves from 250W to 258W and every prescription in the plan moves with it, roughly 3% faster, call it 5 seconds per kilometre at 70 kg. You can see the causal chain. That is what a model looks like when someone shows you the working.

And the ChatGPT-built plan? It’s a template with no layer two at all. An LLM writing your 16-week block has no memory of Tuesday, no access to your Garmin files, and no state between sessions. Ask it for week 7 twice and you will get two different weeks, which should tell you something. It is a fast, cheap, surprisingly competent plan author that goes blind the moment you close the tab. That’s fine if you’re the one doing the adapting. It is not a coach.

Three questions to ask any app

QuestionTemplate answerModel answer
Does next week change if I change nothing but my running?No, only if I edit settings or press “adjust plan”Yes, silently, within a day of syncing
Can I see the number it’s using and watch it move?A green tick, a streak, a completion percentageCTL/ATL, load ratio, CP, threshold pace, with history
What happens when I beat the prescribed pace for three weeks?NothingPaces re-baseline faster

That third one catches more apps than you’d think. Genuine adaptation runs in both directions. If you’re consistently holding 4:25/km on sessions prescribed at 4:35 with your heart rate 8 beats lower than the equivalent session a month ago, a model re-estimates your threshold and gets harder. A template waits for you to go and run another time trial and retype it.

Run question two as a literal exercise tonight. Open the app and try to find a single number about yourself that has a history you can scroll through. If everything it shows you is about the plan rather than about you, the plan is the only thing it knows.

What to do with the answer

If you’ve got a template, you haven’t wasted your money. Runna and the Garmin Coach plans are better than what most recreational runners write for themselves, and structure beats improvisation for almost everybody. You just have to supply the adaptation manually. That means three habits: keep your own weekly volume log and don’t let any week exceed the previous four-week average by more than about 10%, put a genuine down week (60 to 70% of volume) every third or fourth week even if the plan doesn’t, and check that quality work sits around 10 to 20% of weekly kilometres rather than the 30% a template can drift into when you cut days per week but keep the sessions.

If you want the app doing that arithmetic instead, you are shopping in a much smaller market, and it’s worth reading the comparison work rather than the landing pages. Our tested breakdown of AI coaching apps goes through which ones actually recompute and which ones just re-render.

One more thing about the word itself. “AI” in this category almost never means a neural network predicting your race time. It means either a rules engine, a 50-year-old differential equation, or an LLM writing text. Those are three genuinely different products wearing one label, and the label is doing a lot of work that the software isn’t.