#.Overview
Tandemise is a local-first desktop app for running software work with a team of AI agents instead of a single chat window. You describe an outcome in one sentence and say what “done” means. Tandemise plans the work across specialized roles — product, design, architecture, development, review, QA, release — runs each task in its own Git worktree on whichever agent you’ve assigned to that role, collects the typed documents each role hands to the next, checks gates measured from facts, and only interrupts you for decisions that deserve a person. It stops at a verified release candidate, and the call to ship is yours.
It isn’t a model, a coding agent, or a chat client. It’s the organization layer above them. The test the codebase is built to pass is one sentence in its README: if Claude Code disappeared tomorrow, the organization should survive. Missions, roles, decisions, artifacts, permissions, and history stay intact; only an adapter changes.
Tandemise is open source under the Apache 2.0 license, at version 0.8, and very much a work in progress. Like Beveloce, it’s a project I build to learn — in this case, how to make agents dependable enough to hand real work to.
#.Origins
On January 30, 2026, OpenClaw launched, and like a lot of developers I spent the following days trying it. It was a genuinely exciting experience: an agent that could keep working on its own, with real tools. It also made something clear to me. A capable agent with access to everything is great to talk to and hard to trust with a codebase. What was missing wasn’t a smarter agent — it was the structure a team has around its people: who does what, what gets handed over, who signs off, and how you know something is actually done.
In February 2026 I built a first version of that idea for myself: a small harness where agents took on roles and handed each other documents instead of chat transcripts. When I started Beveloce in March, that harness grew inside it as Ergo Sum — very capable, but tied to Beveloce’s own tooling and to how that one product is planned, built, and verified.
By July 2026 it was clear the interesting parts had nothing to do with cycling. Roles, typed handoffs, gates, approvals, and budgets would work for any codebase. So I ported most of the code into a new repository, stripped out everything specific to Beveloce, and gave it a name of its own: Tandemise — people and agents working in tandem. Later I open-sourced it on GitHub, and the latest release, 0.8.0, added evals.
#.The Idea
The design rests on a short maxim: Tandemise owns the organization. Agents are replaceable workers. Tools are replaceable capabilities. Machines are replaceable execution targets.
That leads to a few decisions that shape everything else:
- Adapters, not integrations. Every agent runtime implements one contract. The scheduler routes by capability — who can do shell and git, is healthy, and isn’t saturated? — never by vendor.
- Artifacts, not chat. Roles hand off typed documents with schemas: a product spec, an architecture plan, a change set, a review report, a QA report. A reviewer reads the diff and the spec, never the implementer’s transcript, because independence is the whole point of a review.
- Gates, not assurances. Progress depends on measured facts, like
checks.typecheck == PASS && review.blocking_findings == 0. “The developer says it’s done” is never a gate condition. When a gate blocks, it tells you exactly which condition failed and what value it measured. - Default deny. No tool, path, or domain is reachable without an explicit grant scoped to a single assignment. A QA agent can’t even discover a tool that writes to production.
#.The Experience
All the screenshots here are real captures of the desktop app, taken against a demo workspace — a made-up product called Taskly, run by a scripted demo agent — so nothing on screen is anyone’s actual work.
#.A mission, planned across roles
A mission starts as a sentence — here, “Let people download any report as a CSV file they can open in a spreadsheet.” Tandemise turns it into a plan: a product manager writes the spec, a QA engineer verifies it, a release manager assembles the candidate. Each card shows who did the work, who is responsible for it, the branch it ran on, and whether its gate passed. When the release gate refuses to pass because two criteria are still unverified, the plan says so in plain words.
#.“Done when,” verified line by line
A mission can’t be planned until it says what done means. If you haven’t, a product agent proposes criteria and asks only the questions that would change the plan. Every line becomes a numbered criterion, the spec has to cover each one, and QA has to verify each one by its ID. The mission shows the result as a checklist — and when something can’t pass, the decision card spells out your options: leave it blocked, accept the result and continue, or retry once more.
The gates behind that checklist are visible too. Each one shows its condition and the value it measured, so “blocked” is never a mystery.
#.An inbox for what needs a person
Agents shouldn’t ask permission for everything, and they shouldn’t silently get stuck either. The inbox holds only what truly needs you: plans waiting for approval, results that ran out of retries, a mission that hit its spending limit, a refinement with open questions. A mission that nothing is moving gets a single Stalled row with the one action that would move it; an agent that has gone quiet gets a Quiet row with Keep waiting or Stop and retry.
Around the engine sits the loop a product owner runs every day: rank requests in a backlog with a limit on how much runs at once, cap agent minutes, tokens, or dollars per mission and per month, and put standing work — weekly dependency updates, a Friday report — on a schedule. The home screen counts what needs you, what’s in progress, what’s been verified, what it cost, and what’s stuck. The status report writes the same picture from stored facts, with no model involved.
#.A team, with owners
Every agent on the team has a human owner. Roles are first-class: each one declares the documents it produces and consumes, the capabilities it’s granted by default, the quality bar an evaluator checks its output against, and the model it runs on.
#.Models, Skills, and Evals
Tandemise drives agent runtimes rather than calling model APIs directly. Claude Code is detected automatically, Codex has its own adapter, and any other command-line agent — Gemini CLI, opencode, cursor-agent — can be wired in through a generic adapter without writing code. Model names pass straight through to the runtime, so there’s no list to keep up to date.
That makes model choice a property of the organization, not of a chat session. Each role, and each step of a workflow, can run on the model that fits it. A retry can escalate to a stronger model, an economy model takes over as a mission nears its spending limit, and a review gate can require a different runtime or model from the one that wrote the code, so a model never grades its own homework. Metrics show usage by model: runs, agent time, tokens, and cost.
You can bring the skills you already use — from ~/.claude/skills, a folder, or a Git repository — and pin a specific version to the roles that need it. Every run gets exactly the version it was pinned to, which makes a skill change something you can compare rather than something that quietly shifts behavior.
Which brings me to evals, the part I’ve enjoyed building most. Every agent step with a completion gate is scored automatically on every attempt: whether the gate passed, which criteria it left verified, failed, or unverified, tokens, cost, wall time, and attempts. A From your runs view reads those scores back by role and model — first-attempt pass rate, attempts to pass, median cost, and median time. And any finished step can be saved as a case, frozen with its inputs, starting commit, and mission context, then replayed with a different model, a different skill version, or a whole different setup, with the results side by side.
Two rules keep this honest. Scores are measured facts, never judgments, and the engine that runs your missions never reads them back — a good or bad score doesn’t change how the next run behaves. And eval trials never mix with your project’s own history, so trying a candidate can’t quietly move your real numbers.
#.MCP and Guardrails
Tandemise uses MCP in both directions. Toward the agents, it runs an MCP gateway that exposes to each worker only the tools it has been granted for that assignment. I verified this with a real Claude Code session started in strict MCP mode, which saw only its granted tools and nothing else. Toward the outside world, it connects to hosted MCP servers such as Figma, Canva, Linear, Notion, Jira and Confluence, Vercel, and Sentry through OAuth 2.1 with PKCE, and it can also run local MCP servers.
The rest of the safety model follows the same deny-by-default instinct. Secrets live in the macOS Keychain, and the database only stores references to them. Environment variables are checked so secrets don’t leak into agent runs. The autonomy defaults are conservative: plans ask for approval, local code changes run on their own, external writes follow policy, production releases ask, and financial actions are denied.
#.Under the Hood
A daemon, tandemd, is the single authoritative process. It exposes an authenticated HTTP and WebSocket API on localhost and is the only place in the codebase that names a concrete provider. The Electron and React desktop app is a control surface only, with no access to the filesystem, shell, or database. A small Swift helper handles the macOS pieces — accessibility, screen capture, and input.
The monorepo is TypeScript on Node 22, split into 21 packages in strict layers: domain, application, and kernel at the core; runtimes, execution, persistence (SQLite), a content-addressed artifact store, integrations, and browser and desktop control on the outside. The dependency rule is enforced mechanically — a build check fails if a package imports from an equal or higher layer, or if a core package imports something like Electron, Playwright, SQLite bindings, or child processes.
Crash recovery is treated as a feature. The daemon outlives the window, runs are checkpointed with resumable sessions, and a restart reconciles interrupted work and records in the mission timeline that it happened.
#.How I Built It
Tandemise is built the way it asks you to build: with agents, in phases, with evidence. Each phase has a written spec, a plan, and screenshot evidence of the result. Architecture decisions are recorded as ADRs. The MVP has 21 acceptance criteria, each marked as met only with the script that proves it. CI builds everything, enforces the layer boundaries, the design rules, and dependency licenses, type-checks the desktop app, and runs the offline end-to-end checks. Pull request titles follow Conventional Commits, and release-please cuts the versions. Security-sensitive work gets two independent reviewers, run adversarially against the code.
#.Inspiration and Credits
If all of this makes it sound like I’m somehow smart, I should set the record straight: I’m not. I’m just curious. Most of these ideas are borrowed from other products and open-source projects. So why build it myself? Because that’s how I usually learn something new.
Tandemise is heavily inspired by a handful of projects working out how people and agents get real work done together, and most of its shape comes from them rather than from me: the phases a mission moves through, how work is split across roles and shared between people and agents, and an operating model of budgets, approvals, and standing work. OpenClaw sparked the idea; these are the projects the design draws on most:
- Paperclip showed that the interesting layer sits above the agents: an organization with roles, goals, budgets, and governance that runs whatever agents you already have. The organization layer, the roles, and the spend and time limits come from there.
- Grok Bot treats agents as persistent teammates that keep working on their own and come back only when something needs your approval. Tandemise’s approvals, its inbox, and its routines that queue standing work on a schedule follow that model.
- Buzz puts people and agents in the same workspace as equals, next to the code, without locking you into one vendor’s models. Tandemise borrows its owned-agent model (every agent has a person who owns it and answers for its work), owner-reviewed drafts, human steps in workflows, and runtimes as replaceable adapters.
- Hermes Agent is an agent that grows with you: it builds skills from experience and remembers what it has done. The idea that missions, decisions, and history should outlive any single agent or session comes from there.
What Tandemise adds is mostly the assembly: putting these ideas together in one local-first place, with typed artifacts and gates measured from facts. Thank you to the people behind each of them.
#.Where It Is Now
Tandemise is at version 0.8, runs on macOS, and is open source at github.com/demogar/tandemise, with a site at tandemise.com. It’s early, and it says so on the tin. There’s no user base I’m claiming. It’s a working system I keep improving because every phase teaches me something about building with agents that I couldn’t have learned from reading about it.
#.Reflections
The organization matters more than the agent. Models get better every few months; the structure around them — roles, handoffs, gates, owners — is what makes their work usable. Designing for replaceable agents turned out to be a very freeing constraint.
Facts, not vibes. The single rule I’d keep if I had to drop everything else: no model output decides whether work is done, planned, stopped, or stuck. The daemon decides from stored facts and fixed rules. That’s what makes it possible to walk away and trust what you come back to.
Ask people less, but ask them better. An agent system that asks for approval constantly gets rubber-stamped; one that never asks gets ignored until it breaks something. The inbox, the Stalled and Quiet rows, and the decision cards are all attempts to find the middle: one clear question, with its consequences spelled out, at the right time.
You can’t improve what you don’t measure. In Beveloce, a verifier agent categorizes how each model’s output went on real work. Evals, scored from gates and replayable on another model, are how I want to finish that thought.
