#.Overview
Beveloce is an adaptive cycling coach that lives on your phone and your laptop. It reads your real training and health data, combines it with the goal you’re working toward, prescribes a plan week by week, looks at how you actually rode it, and adjusts what comes next.
I’ve been building it since March 2026, and it’s still very much a work in progress. It’s also, first and foremost, my way of learning AI engineering in a professional context: putting language models inside a real product, wiring agents and MCP servers into my own development workflow, comparing models from different providers on the same work, and building the harnesses and guardrails that make any of that trustworthy.
I want to be clear about what it isn’t. Beveloce doesn’t replace a coach, and that was never the intention. A good human coach reads things no dataset holds: the way you talk about a hard week, what’s happening at home, the difference between tired and burnt out. What I wanted to understand is narrower and, to me, more interesting as an engineer: how far can you get by combining real health data with a training goal, prescribing a plan, watching how that plan gets executed, and adapting it along the way — and where exactly should a model be trusted in that loop, and where shouldn’t it?
#.Why a Cycling Coach
I could have learned all of this with a to-do app and a chatbot. I chose a domain I know and care about — I ride — because it has the properties that make AI work hard in the real world:
- The data is real and messy. Power files, heart rate, HRV, sleep, resting heart rate, and subjective check-ins arrive from different services, at different times, with gaps.
- The output has consequences. A plan that ramps too fast, or piles intensity on intensity, isn’t a typo — it’s a bad month for the person riding it.
- It’s a loop, not a prompt. You prescribe, the world happens, and you have to reconcile what was planned with what was done before you prescribe again. That’s the shape of most agentic systems, in miniature.
- There’s an established body of knowledge. Periodization, training load, and intensity distribution are well studied, so there’s something solid to check a model’s output against instead of taking it on faith.
That combination made it a much better teacher than a toy project would have been.
#.The Loop
Everything in Beveloce serves one loop. It starts from real data, turns it into a picture of where you are, drafts a week toward your goal, and checks that draft with deterministic code before you ever see it. After you ride, it reads back what actually happened and adapts the next week from there.
The banner on that diagram is the most important design decision in the project: the model proposes, code decides the numbers, and you confirm the changes. I’ll come back to why.
#.The Experience
The mobile app has four tabs — Today, Timeline, Progress, and Profile. The screenshots below are real captures from the app, running against a seeded demo account so no one’s personal data is on screen.
Today answers one question: what am I riding, and why? Each session is a real structured workout — warm-up, the work, cool-down — with duration, load, a power target derived from your FTP, the zone it lives in, and a one-line note from the coach explaining its place in the week. Timeline shows the whole block: where you are in the build, when the peak starts, and the load of every week along the way. When a week is done, it shows what was planned next to what happened.
Adapting is where it gets interesting. Weeks carry a check-in — felt about right, felt too hard — and sessions record whether they were completed, skipped, or missed. That’s the signal the next week is built from. Progress turns your history into something readable: a rider profile of strengths and where there’s room to grow, fitness and form over time (CTL, ATL, and TSB), FTP progression next to the estimated FTP from Intervals.icu, and an early read on where your current trend is heading.
Sessions export straight to the places people actually train — Zwift, Rouvy, MyWhoosh, Garmin Connect through .fit files, and Intervals.icu — so the plan isn’t a paragraph of advice; it’s a workout your trainer can execute. And there’s the feature where it all started: give Beveloce a GPX route and it reads the terrain and turns it into an indoor workout that asks for the same effort.
#.What I Set Out to Learn
#.Putting an LLM inside a product
The first version of Beveloce, in March, wasn’t a coach at all — it was a GPX-to-Zwift route converter. The coach arrived in April, with a design spec, the first coaching tables, and a fetcher for fitness data from Intervals.icu. By the end of May, the model sat behind a pluggable provider interface with two tiers:
- a smart tier for the expensive work of drafting a week, and
- a fast tier for chat replies and short coach notes.
Two providers implement that interface: OpenRouter, which lets me swap models from different vendors by changing configuration rather than code (and falls back to the other tier’s model if one fails), and Anthropic’s API directly, where the long system prompt is cached so repeated generations don’t pay for it every time.
Out of the box, neither tier is pinned to a single model. Generations rotate round-robin across different models, and a separate agent verifies each output and categorizes how it went. Over time, that turns every real generation into a data point about how one model behaves compared with the others, on the same job.
Running a model in a product also means deciding what it’s allowed to cost. Plan generations and chat turns are capped per person per day, high-volume work goes to the cheap tier, and the rider’s language — the web app speaks English, Spanish, French, and Italian — is applied in one central place, so no prompt can forget it.
#.The model proposes; code decides
This is the lesson I’d most want another engineer to take from the project.
Early on, it’s tempting to let the model write the whole week: sessions, durations, targets, all of it. It reads well. The problem is that a fluent answer isn’t the same as a correct one, and training plans have hard rules — how much load can rise from one week to the next, how intensity should be distributed, what a recovery week actually looks like, how to taper into an event.
So the model’s draft goes through a set of deterministic guardrails before anyone sees it. Volume is fitted to the week’s time budget. Intensity is fitted to a target distribution with a floor on easy riding. Week-over-week load is capped. Deload weeks and missed weeks have their own rules. If a draft can’t be repaired, the model gets exactly one corrective retry — not an open-ended loop. And the numbers you see on a session — duration, load, zone — are computed by code from the workout’s structure, not copied from what the model said about it.
A comment at the top of the re-planning service says it better than I can: the LLM proposal is a hope; the guardrails are the check. The coach chat follows the same rule. It never writes workouts itself; it translates “make this week easier” or “I only have three days” into a bounded adjustment that goes through the same checks. The principles behind those rules are tied to published training research — Seiler and Kjerland on intensity distribution, Issurin on periodization, Coggan and Allen on FTP-anchored power zones — so the coach can say why, not just what.
#.Adapting to what actually happened
A plan is only as good as its second week. Beveloce reads completed activities back from Strava and Intervals.icu, matches them to the sessions that were planned, and grades how closely each one was executed. Wellness data — HRV, resting heart rate, and sleep from Oura, Apple Health, or Health Connect — joins the subjective check-ins.
From there, a router decides whether a change deserves a small reshaping of the current week or a new plan altogether. One-tap presets cover the common realities (“I’m sick,” “life happened”), a minimum-dose shape keeps a week useful when time collapses, and a Monday review proposes adaptations that you confirm before they’re applied. The model can suggest; it doesn’t get to silently rewrite your month.
#.Agents, harnesses, and Ergo Sum
The other half of the learning happened in how Beveloce gets built. Most of its code is written with AI agents, and making that productive — rather than a faster way to generate rework — became a project of its own.
Early on, Beveloce’s workflow ran on Ergo Sum, a harness I built for it: a set of repeatable procedures and specialized agents tied closely to how this product is planned, built, and verified. Today that workflow lives in a Claude Code plugin shared by the web and mobile repositories. It has skills for starting a story from an issue, building a feature (plan, design, test-first implementation, gates, review, pull request), fixing a bug with the reproduction written first, verifying acceptance criteria, reviewing pull requests, splitting a cross-surface feature into an epic, and discovering and scoring new ideas. It also has debates: three product-manager agents, each on a different model, argue about improvements to shipped features while a neutral synthesist writes up the result.
Behind the skills sit agents with deliberately different jobs and model tiers — planners on the strongest reasoning models, an implementer on a solid coding model, and exploration and commit-pushing on the cheapest models that can do them well. Each repository declares its own quality gates in a small profile file, so the same workflow knows what “done” means for the web app versus the mobile app.
The generic parts of Ergo Sum — roles, typed handoffs, gates, approvals — outgrew Beveloce. In July 2026 I moved them into their own repository, and that became Tandemise.
#.Comparing models from different providers
Once the workflow was a harness with defined roles, it became a place to compare models honestly — not in a chat window, but on the same real work, measured by cost and by whether the output passed its gates.
The routing moved a lot over the summer. It started on Claude (Opus for planning, Sonnet for implementation, Haiku for cheap reads) and grew an OpenCode adapter that mirrors the same agents and commands so the harness could run on other vendors’ models. From there it moved to OpenAI’s GPT models, then to cost-tiered routing across providers: DeepSeek for cheap, bounded work, GLM as the primary model for a stretch, and a self-hosted gateway in front of DeepSeek as a fallback. Today the table reads: GPT 5.6 for discovery, plans, reviews, and orchestration; Kimi K2.7 Code for implementation steps; Qwen 3.7 Plus for verification against acceptance criteria, because it can read screenshots; DeepSeek V4 Pro for bulk exploration and small changes; and DeepSeek V4 Flash for read-only status and commits. The Claude routing stays available as a mirror.
The most useful finding wasn’t about which model is “best.” It was about where the cost actually is. In a long orchestration session, the whole context is re-sent on every turn, so the orchestrator dominates the bill. When that loop ran on a metered model for a couple of days, it burned through a meaningful share of the month’s budget. The fix was structural: the long session runs on a flat subscription, and metered models only ever see small, bounded dispatches with a clear objective. That’s the kind of thing you only learn by measuring.
#.MCP in the daily loop
MCP servers are part of the everyday workflow rather than a demo. An Intervals.icu MCP server, for example, let me explore real training data — activities, wellness, fitness curves — from inside the agent session while I was designing what the coach should read and how. Tandemise later took this further, with an MCP gateway that serves each agent only the tools it has been granted.
#.Evals: from verification to replay
The coach already grades its own models. Because generations rotate across models and a verifier agent categorizes every output, each model builds up a record of how it did on real weeks, side by side with the others. Underneath that sits the deterministic safety net: the guardrails, a large suite of unit tests around them, and end-to-end runs against a stubbed model, so the product’s behavior can be tested without calling a real one. On the build side, models are compared by gates passed, attempts, and cost.
What comes next is making those comparisons repeatable: a fixed set of athletes and weeks, replayed on every model and every prompt change and scored the same way, so swapping a model becomes a measured decision rather than a hunch. The evals I built into Tandemise work exactly like that: save a finished step as a case, replay it on another model, and compare the scores side by side. They’re the prototype of that idea.
#.The Stack
- Web: Next.js and React with TypeScript and Tailwind, on Supabase (Postgres and auth).
- Mobile: React Native, reading Apple Health and Health Connect with read-only, least-privilege permissions.
- Mobile CI/CD on EAS: merging a release-please pull request publishes a release. EAS Build then leaves an installable preview build behind it, and EAS Update ships JavaScript-only changes over the air to each release channel. Shipping to the stores is a deliberate, manual step, never a side effect of a merge.
- Shared code: packages for the domain logic, the design system, and a typed API client, with a public OpenAPI contract checked in CI.
- Observability: product analytics on PostHog, errors on Sentry.
- Testing: test-first throughout, with Vitest and Playwright (including automated accessibility checks) on the web, and Jest and Maestro on mobile.
- Workflow: conventional commits, and release-please for versioning.
Across the three repositories — web, mobile, and the agent tooling — that adds up to roughly 1,850 commits since March. That number isn’t about traction. It reflects how much of the learning happened in small, reviewable steps.
#.Where It Is Now
Beveloce runs on the web at beveloce.com in four languages, and the mobile app is in progress. I built it for myself and a group of friends, mostly in the United States and Europe. We’re all amateur cyclists who ride in different styles, and we use it for fun. Nothing more. A Paddle checkout is connected, but it’s mainly there so friends can donate to the cause if they want to.
It’s a small project and I treat it as one: there’s no audience I’m optimizing for, and a lot of the roadmap exists because I want to learn the thing it requires. The next item is the coach evals described above.
#.Reflections
A few things I believe more firmly now than I did in March.
Trust has to be designed, not hoped for. The most valuable code in Beveloce isn’t the prompt; it’s the deterministic layer that decides whether the model’s answer is allowed through. Most “AI features” live or die on that boundary.
Models are interchangeable; roles and measurements aren’t. Once the work was split into roles with clear inputs, outputs, and gates, switching models became a configuration change — and comparing them became possible at all. That idea turned out to be big enough to become its own project.
Cost is an architecture decision. Where the context lives, which step runs on which tier, and what gets cached matter far more than per-token prices.
Real data keeps you honest. Plans that looked perfect on paper met sickness, travel, skipped sessions, and bad weeks. Building for the second week, not the first, is what made this a loop rather than a generator.
Beveloce isn’t finished, and I don’t expect it to be for a while. It’s the place where I practice the parts of AI engineering that matter in a professional setting — integration, guardrails, agents, cost, and measurement — on a problem I see as difficult and interesting.
