AI Coaching Apps, Tested Against Each Other
The interesting question about AI running coaches stopped being “can it write a plan” about three years ago. Any of them can write a plan. The question now is narrower and harder: when your body, your calendar and your last four weeks of training disagree with the plan, which app notices, and what does it do about it?
I’ve been running the same set of tests across the main contenders since early 2025. Same athlete profile, same interruptions, same deliberately awkward inputs. What follows is what separates them when you push, plus how to audit whatever you’re already using without switching.
The test setup, so you can judge the results
Every app got the same synthetic athlete: 41 years old, 5K PB of 21:30, half marathon 1:42, running 4 days and roughly 38 km a week for the past three months, training for a spring marathon 18 weeks out with a target of sub-3:45. Where the app accepted Strava or Garmin history, I connected an account with 14 months of real data behind it (my own, with the paces scaled).
Then I ran four scenarios against each plan:
- The missed week. Log nothing for seven days, then open the app on day eight. A chest infection, in the story.
- The too-fast easy runs. Execute every easy run 25-30 seconds per km faster than prescribed, at 82-85% of max HR instead of the 70-75% the plan implies.
- The failed session. Hit the first three reps of a 6 x 1K at target, then fade 12, 18 and 26 seconds off pace on the last three.
- The compressed block. Tell the app mid-plan that the race moved four weeks earlier.
Those four cover most of what actually happens to a recreational marathon build. The plan that survives contact with them is the plan worth paying for.
What the apps did with the missed week
This is the cleanest separator, and the results were not close.
Runna handled it best of the subscription apps. Reopening after seven blank days, it offered to reschedule and pushed the remaining plan back, which cost me one week of the build and shortened the taper by nothing. Critically, it did not try to make up the missed mileage. Week 9 came back at 41 km against the 46 km it had originally held, a deliberate step down before resuming the progression. That’s the right instinct: after a week off sick, you re-enter below where you left, not at it.
Garmin Coach (the Greg McMillan marathon plan on a Forerunner 265) simply carried on. Week 9 arrived as originally written, 48 km with a 2:15 long run, no acknowledgement that the previous week was empty. Garmin’s daily suggested workouts do respond to training status and HRV via Body Battery and acute load, so a downstream correction happens, but the plan structure itself is static. If you follow the plan rather than the suggestions, you get no adjustment at all. The detailed head-to-head is in Runna vs Garmin Coach: A Side-by-Side Test, including what happens when the two disagree on the same day.
TrainingPeaks with the AI plan builder rebuilt properly but overcorrected, dropping the next three weeks to 32, 35 and 38 km and delaying the first 30 km long run by a fortnight. Safe, but on an 18-week build that costs you a genuine chunk of specific endurance.
A ChatGPT-built plan (GPT-5, given the full athlete profile and asked for an 18-week marathon plan as a table) did whatever I asked it to do, which is exactly the problem. Prompted with “I missed a week with a chest infection, what now,” it gave sensible generic advice: return at 60-70% volume, no quality for the first few runs, build back over 10-14 days. Prompted with “I missed a week, how do I catch up,” it cheerfully compressed two weeks of work into one and produced a week with two long runs in it. The model has no memory of your load and no stake in your outcome, so it mirrors the framing of your question. That’s a structural weakness, not a prompt problem.
The tell to look for in any app: after a break, does next week’s volume come back below the last completed week? If it comes back at or above, the app is following a calendar, not you.
Easy pace drift, and who actually catches it
Scenario two is the one that quietly wrecks marathon builds. You run your easy runs too hard, your legs are never fresh for the sessions, the sessions degrade, and eight weeks later you’re flat.
Here’s what the drift looked like in the data. Prescribed easy pace was 5:55-6:15/km. Executed:
Week Prescribed Executed Avg HR %HRmax TRIMP
1 6:05/km 5:38/km 152 82% 118
2 6:05/km 5:36/km 154 83% 126
3 6:00/km 5:33/km 156 84% 134
4 6:00/km 5:31/km 158 85% 141
Four weeks of that, with weekly load climbing 6% purely from intensity rather than volume. Every app had the HR data.
Runna flagged it in week 3 with a note on the easy run card suggesting I slow down, and it adjusted the prescribed range downward, which is a slightly odd response (the ranges were already correct; I was ignoring them). It did not change the session load elsewhere to compensate.
Garmin was the most informative here, though not through the plan. Training Status flipped to “Unproductive” in week 3 and the Load Focus chart showed low aerobic at 31 points against a target range of 180-560, high aerobic well over the top of its range. That is a genuinely accurate diagnosis presented in a form most runners never open. The Forerunner also started suggesting shorter easy runs via daily suggested workouts, implicitly compensating.
TrainingPeaks showed it on the PMC (a TSB of -28 by week 4 with CTL barely moving) but said nothing. You have to read the chart yourself.
ChatGPT and Claude, given the table above pasted in as text, both identified the problem immediately and correctly, and both wrote a better prescription than the apps did: cap easy runs at a specific HR ceiling of 145, walk hills if needed, and re-test in two weeks. General-purpose models are strong at diagnosis when handed clean data and weak at noticing without being asked. That split is worth internalising, because it tells you how to use them.
The failed session tells you the most about plan quality
Fading 26 seconds off pace on the last rep of a 6 x 1K is a meaningful signal. It usually means the target was wrong, the athlete was underrecovered, or both. What the app does next reveals whether there’s any model of the athlete underneath.
Runna’s response was to keep the following week’s interval target identical. It treats prescribed paces as derived from your race times and recent efforts, updating them when you log a race or a time trial, not when you fail a session. So a session failure is invisible to it unless you manually adjust your fitness inputs.
Garmin, running a structured workout from the plan, logged the workout as completed and recalculated a race predictor slightly downward. The next quality session came from the same plan template at the same relative intensity. Its adaptive layer sits in daily suggested workouts, which are excellent (they’re built on your acute:chronic load ratio and HRV status), but plan workouts override them, so the adaptivity is switched off precisely when you’re following a plan.
The only tools that reliably caught the failed session were the ones with an explicit fitness model. Intervals.icu, which is free and which most runners have never opened, computes a power/pace duration curve and flagged that my 1K efforts fell below the modelled curve, a straightforward way to see that the target was too aggressive. Feed the same six splits into a chat model and it will also spot it, with more nuance about why:
Rep Target Actual Delta
1 4:15 4:14 -1
2 4:15 4:16 +1
3 4:15 4:17 +2
4 4:15 4:27 +12
5 4:15 4:33 +18
6 4:15 4:41 +26
That shape, three clean reps then a cliff, is glycogen and neuromuscular fatigue rather than aerobic incapacity. A progressive fade from rep 1 would mean the target was simply too fast. An app that treats both patterns the same is not modelling you.
Moving the race date exposes the planning engine
Telling an app the race is now four weeks earlier is a stress test of its internal logic. Marathon plans have a required structure: a peak long run of 28-32 km, at least two or three of them, a taper of 2-3 weeks, and enough marathon-pace work to calibrate.
Runna rebuilt the plan in a few seconds and produced something defensible, with the peak long run moved forward and the taper preserved at two weeks. It dropped from three long runs over 28 km to two. Reasonable call.
Garmin Coach rebuilt too, and its output was more conservative: it shortened the build rather than the taper, which on a compressed timeline costs you the specific endurance you most need. It also kept the same number of weekly workouts, meaning weekly load jumped by about 9% to fit the same progression into fewer weeks. That’s the wrong lever to pull.
The chat models, when asked to compress, split. Given no constraints, they compressed everything proportionally, including the taper, which is the one part that shouldn’t move. Given the constraint “keep the taper at 14 days and the peak long run at 30 km,” they produced good plans. The lesson repeats: these tools are excellent executors of well-specified constraints and unreliable at supplying the constraints themselves.
How to audit whatever you’re already running
You don’t need to switch apps. You need four numbers, and any of these tools will surface them if you know where to look.
Weekly volume progression. Chart the last eight weeks. Week-on-week increases should mostly sit in the 5-10% range, with a down week every third or fourth week dropping 25-35%. A plan that goes 42, 46, 50, 54, 58 with no down week is not progressing you, it’s queueing an injury. Strava’s own weekly bar chart does this fine; so does the Fitness page if you have a subscription.
Intensity distribution. Over four weeks, add up time (not distance) by HR zone. You want roughly 75-85% of total running time below your aerobic threshold. Garmin gives you this directly in Load Focus. In Strava, sum your zone times manually across a month; it takes ten minutes and it’s the single most useful audit here.
Long run as a share of the week. Should be 25-35% of weekly volume. If your long run is 32 km in a 60 km week, that’s 53%, and the other four days aren’t supporting it. This is the most common failure mode in AI-generated marathon plans, and both the chat-built plans and one of the app plans did it at least once.
Acute:chronic ratio. 7-day load over 28-day load. Below 0.8 you’re detraining, above 1.5 you’re accumulating risk. Intervals.icu shows this free after a Strava connection; Garmin approximates it in Acute Load; TrainingPeaks does it as TSB.
Here’s the audit run against a real 4-week window from one of the test plans, laid out as you’d want to see it:
Metric Value Healthy range Verdict
Weekly volume change +4/+9/+8/-29% 5-10%, down wk every 3-4 OK
Easy-run time below AeT 61% 75-85% FAIL
Long run share of week 38% 25-35% Marginal
Acute:chronic (day 28) 1.42 0.8-1.3 Warning
Two failures, both intensity-related, neither of which the app itself raised. The plan wasn’t badly written. The execution drifted, and the app had no mechanism for noticing.
Where each tool actually earns its place
After all four scenarios, my honest read on the current field:
Runna (about £16/month) is the best structured-plan generator for most recreational runners. Its sessions are well-constructed, its paces are sensibly derived, and it reschedules gracefully around life. Its weakness is that it’s a planner with adaptive edges rather than a coach: it responds to what you tell it (a new race time, a missed run) more than to what your data shows.
Garmin Coach is free with the watch and its daily suggested workouts are the best genuinely-adaptive engine in consumer running, driven by HRV status, sleep and acute load. The problem is architectural: when you follow a structured plan, that engine takes a back seat. The best Garmin setup for many runners is actually no plan at all, just daily suggestions plus your own weekly structure.
TrainingPeaks gives you the best diagnostic charts and the worst experience of being told anything. It’s an instrument panel, not a coach. Excellent if you’re willing to read it.
Intervals.icu costs nothing and does more analysis than anything on this list. Connect Strava, give it a week, and look at the fitness chart and the pace-duration curve.
Chat models (ChatGPT, Claude) are the strongest analysts on this list and the weakest monitors. Paste in four weeks of splits, HR averages and weekly totals and ask “what’s the biggest problem here” and you’ll get a better answer than any app’s automated flag. But they won’t ask you for that data, and a plan they write in one conversation has no connection to what you actually run afterward.
A practical setup that works
The configuration I’d suggest for someone already running an AI plan: keep the plan you have. Add Intervals.icu for free and check the fitness chart once a week, Sunday evening, for two minutes. Once a month, export or screenshot four weeks of weekly volume, zone distribution and long-run shares, paste it into a chat model, and ask it to tell you what’s wrong rather than whether things look fine. The framing matters: “is this okay” gets you reassurance, “what’s the biggest risk in this block” gets you analysis.
And before any key session, look at one thing: what did your last three easy runs look like relative to their prescription? If they crept 20 seconds per km fast, the session is already compromised. No app will tell you that in the moment. You will, if you look.
In this section
The supporting pages under this subject.