[ RELAY // HARNESS ]
SYS.ACTIVE
Open-source terminal coding agent

Any model.
One harness.

Relay reads, runs, edits and verifies code from your terminal, and wraps every model in the same four rules: alignment, evidence, context and skills.

Built on pi by Mario Zechner & Earendil Works

4
Harness pillars
0
Extra model calls
13
Open packages
MIT
Free forever
[ 01 // WHAT IS RELAY ]

A coding agent with a harness around it

Built on pi, the agent works through multi-step tasks with a hosted provider, a subscription or a local endpoint. The harness is the part that stays the same when the model changes, and it needs no extra model call.

Alignment

Keeps your standing rules as state, asks before destructive or external commands, and won't edit when you only asked a question.

Evidence

Records what changed and which checks passed. A "done" without a passing check after the last change gets a verification request.

Context

Recent tool results go verbatim, stale ones shrink to one-line stubs, and a compact progress digest keeps the thread.

Skills

Treats SKILL.md as a router, measures which resources lead to actions, and audits how each skill is structured.

Laya router

laya/auto picks the model, effort, role and validation for each request: the cheapest combination likely to succeed.

Any provider

Claude, GPT, Gemini and more through a subscription or API key, custom providers, or a local llama.cpp server.

[ PREVIEW // RELAY IN YOUR TERMINAL ]
$ npm install -g --ignore-scripts @relay-harness/coding-agent $ relay relay · escape interrupt · / commands · ! bash › /model laya/auto › Fix the race condition in workers laya → strong tier · claude-opus-5-5 · high read src/workers/pool.ts edit src/workers/pool.ts $ npm test exit 0 harness · evidence verified
[ // HARNESS CORE ]

Four pillars, each one answering a measured failure

The harness core is four rules that wrap the agent loop and apply to every model Relay runs: Claude, GPT, Gemini, Grok or a local model. They act on the loop's hooks and events, not on a provider API, need no extra model call and run in microseconds. Each pillar targets a failure that published research measured in real coding agents.

01 // Alignment
38%

of failures in 20,574 real sessions broke an explicit developer constraint

The problem

Constraints live as prose in an earlier message, so they fade as the context grows or gets compacted. Questions are treated as permission to edit, destructive commands run without asking, and agents corrected course without pushback in only 3% of resolved cases.

What the harness does

Keeps your rules as state outside the transcript and restates them next to your latest message. Asks before destructive or external commands. Blocks the first edit in a turn where you only asked a question, and records your corrections.

Tang et al., 2026 · How Coding Agents Fail Their Users ↗
02 // Evidence
75.8%

of failed runs ended in a confident false success (AppWorld; 35.6% on tau2-bench)

The problem

Agents say "done" when the environment shows the work failed. LLM judges anchor on words like "successfully" and stay below 0.65 AUROC, while structural signals reach 0.85 to 0.95. Where an independent component re-checked the state, false success fell from about 45% to 3%.

What the harness does

Becomes that independent component. It records what tool calls really changed and which checks passed; a completion claim with no passing check after the last change gets one verification request. Masked exit codes such as | tail or || true do not count.

Advani, 2026 · From Confident Closing to Silent Failure ↗
03 // Context
71% → 91.6%

task completion when only the last 5 tool results plus a progress summary were kept

The problem

Every tool result stays in context forever. Old results describe files as they were before later edits, so the model acts on stale state and the transcript grows until it is compacted. The same pruning cut tokens by 62.7% and time by 60.2%; without the summary, agents lost track and stopped early.

What the harness does

Sends recent tool results verbatim, replaces older large ones with one-line stubs that say whether the file changed later, and restates a compact progress digest. Stubs advance in batches, so the prompt cache keeps hitting.

Lodha et al., 2026 · Less Context, Better Agents ↗
04 // Skills
1.18 → 3.85

skill resources used per run when SKILL.md routes to on-demand files

The problem

With the same knowledge, organizing a skill as a short router plus on-demand references changed how agents worked: skill use spread across the run (20.7% to 48.4%) and resources that led to an action rose from 76.0% to 84.6%. The gain depended on the task, and a harness that treats skills as plain documents cannot see any of it.

What the harness does

Tells the model to treat SKILL.md as a router, measures which resources are loaded and whether an action follows, and audits each skill's structure for missing or unreachable files.

SkillJuror, 2026 · Measuring How Agent Skill Organization Changes Runtime Behavior ↗
/harness
Inspect constraints, evidence, context and skills at any time
harnessCore.*
On by default; tune or turn off each pillar in settings
Not a sandbox
A safeguard against mistakes, not a security boundary
Read the harness core docs

Figures as reported by each study and cited in Relay's source code.

[ // HOW YOU RUN IT ]

Two ways to put it to work

01

In your terminal

Run relay in any folder, connect a model with /login, and work interactively. Branch, resume and compact sessions without losing history.

02

Inside your software

Script it with print and JSON modes, drive a separate process over RPC, or embed it with the TypeScript SDK. Extend it with extensions, skills and MCP.

[ // LAYA ROUTER ]

Laya picks the right model for every request

Without a router, a typo fix costs as much quota as a concurrency bug. Laya is a small classifier, a fine-tuned multilingual encoder, that runs on your machine in Docker. Select laya/auto and, before each request, it answers 18 typed questions about the task in one call: type, complexity, scope, risk, the tools it needs and whether it touches security. Routing costs no tokens, because the classification happens on your machine.

01

Assess the task

Laya answers the 18 questions in a single local call. A task you taught it answers from memory, and without Docker, keyword rules step in.

02

Apply safety floors

Fixed rules Laya cannot lower: authentication, payments, vulnerabilities, production migrations and high-risk work go at least to the strong tier, with review required.

03

Map tiers to models

Laya thinks in tiers (fast, balanced, strong, frontier), never in model names. A registry maps each tier to the models you have credentials for, so new models need no retraining.

04

Pick by utility

Each candidate is scored on quality, cost, latency and risk, weighted by your cost profile (economy, balanced, quality, critical) and by how often that model has succeeded on similar tasks.

05

Brief the model

The chosen model gets a short plan: the role to take, relevant skills, the tools it needs and how to validate. It goes after your message, so the prompt cache stays valid.

06

Escalate and learn

Repeated failures move the task one tier up, and a rate limit switches to another provider. Every outcome is logged locally and sharpens the next choice.

[ ROUTING // THREE REQUESTS, THREE CHOICES ]
› /model laya/auto › "Fix the text of this button" laya → fast claude-haiku-4-5 minimal › "Fix the race condition in workers" laya → strong claude-opus-5-5 high › "Fix the login token expiry" laya → strong claude-opus-5-5 medium rule: authentication · review required
[ A // THE ASSESSMENT ]

Where the answer comes from

For each request, Relay uses the first source that can answer.

Memory

A request close enough (cosine similarity of at least 0.6) to a task you taught with /laya learn reuses that task's labels. It works even without Docker.

Laya in Docker

Relay sends the request (up to 2,000 characters) and the 18 questions in a single call, with a 5-second timeout and no retries.

Keyword rules

If the server is unavailable, or failed in the last minute, deterministic English and Portuguese rules answer with a fixed confidence of 0.5. The status line shows (rules).

The assessment's confidence is the lowest of the three answers that drive routing: task type, tier and effort. Below 0.5, the tier goes up by one.

[ B // THE CONTRACT ]

The 18 questions

They are the contract between Relay and the model. Answers use abstract tiers and efforts, never model names, so the registry can follow new models without retraining Laya.

The task
task_typecomplexityscoperiskambiguityreasoning_requirement
How to run it
capability_tierreasoning_effortagentvalidation_level
What it needs
requires_writerequires_shellrequires_testsrequires_webrequires_browserrequires_databaserequires_git
Safety
security_sensitive
[ C // THE DECISION ]

From tier to model

A registry maps each tier to models of the providers you are logged in to. The cards show the first choice per tier; other providers follow in preference order.

fast
claude-haiku-4-5
relative cost 0.10
balanced
claude-sonnet-5-5
relative cost 0.35
strong
claude-opus-5-5
relative cost 0.65
frontier
claude-fable-5-1
relative cost 0.90

Every candidate at or above the minimum tier gets a utility score. P(success) starts at 0.90 when the model's tier equals the required one (+0.04 per tier above, −0.18 per tier below) and moves toward that model's observed success on similar tasks.

utility = quality × P(success) − cost × quota_cost − latency × latency − risk × failure_risk
Cost profileWeightsMinimum P(success)
economycost 50 · quality 30 · latency 200.75
balancedquality 45 · cost 30 · latency 15 · risk 10 (default)0.85
qualityquality 70 · risk 20 · cost 100.90
criticalquality 70 · risk 300.95

Pick a profile with laya.policy, --laya-policy or /laya policy. Replace a tier with laya.models, and tell Laya how much subscription is left with laya.quota: scarce quota makes a provider more expensive.

[ D // THE MODEL ]

A small encoder, not another LLM

mmBERT
multilingual encoder, fine-tuned (mmBERT-base)
1,100
synthetic training requests in English and Portuguese
97.2%
validation accuracy after 4 epochs
~644 MB
of weights, served on the CPU in Docker

These numbers overstate real-world accuracy: the test requests come from the same templates as the training data. That is why the safety floors exist, and why teaching Laya from your own sessions matters.

[ E // INSIDE RELAY ]

How it runs and what happens when it cannot

Runs locally, in Docker

On start, Relay pulls the Laya image (CPU or CUDA, picked for you) and keeps a relay-laya container listening only on 127.0.0.1:8737. No Docker? Keyword rules route instead and everything else keeps working.

When Laya is unavailable

  • ·Docker not installed or not running: keyword rules route, and Relay notifies once.
  • ·Image downloading or model loading: rules answer until the server is ready.
  • ·A call fails or times out: rules answer, and Laya is not asked again for 60 seconds.
  • ·A learned task matches: memory answers, even without Docker.
[ F // THE LEARNING LOOP ]

It learns the work you actually do

The shipped model learned from synthetic requests. /laya learn teaches it from your real sessions.

01

Collect

Lists up to 60 recent requests with the evidence of how each went: model, tools, files changed, test results, outcome and your next message.

02

Label

The agent answers the 18 questions for each task from that evidence, and writes a lesson when the task taught something specific.

03

Remember

From the next request on, similar tasks route from memory and the model receives the lessons, with no wait for training.

04

Train

A short-lived Docker container fine-tunes the routing model for 3 epochs, mixing in older exercises so nothing is forgotten.

05

Test and activate

The new and current models answer the same held-out test. The new one, named local-1, local-2…, takes over unless it scores more than 1% worse.

Commands
/laya statusCost profile, container state, routing model, quota and the last plan
/laya setupPull the image and start or replace the container now
/laya policy <profile>Cost profile for this session
/laya learnLabel this session's tasks and train on them
/laya trainTrain again on every task collected so far
/laya modelsShipped and trained models with their test scores
/laya use <model>Route with another model, such as v1 or local-2
/laya stopStop the container
[ // AUTHORED SKILLS, AGENTS & INTAKE ]

Skills, agents and a command written for Relay itself

Besides the core, the repository ships its own skills and agents, plain Markdown files with a short front-matter header, plus the built-in /intake command. They are working examples of how to extend Relay with the same mechanisms you can use.

Skills

Instructions the agent loads only when a task needs them. They live in .relay/skills.

  • release

    Prepare, publish, verify and recover releases with lockstep versioning.

  • security-review

    Review code that runs commands, resolves paths, loads extensions or MCP servers, stores credentials or authorizes tool calls.

  • dependency-review

    Review dependency and lockfile changes, and say what an audit report must contain to count as evidence.

  • test-evidence

    Decide whether a test result is real evidence, for example by seeing a regression test fail before the fix.

  • interactive-testing

    Test and debug the interactive terminal UI in a controlled tmux session.

  • add-llm-provider

    Checklist for adding a new LLM provider, from core types to the test matrix and docs.

Agents

Specialist sub-agents with their own context window, tools and model. They ship with the subagent example extension.

  • scout

    Fast codebase recon that returns compressed context for other agents.

  • planner

    Turns context and requirements into a clear implementation plan.

  • worker

    General-purpose agent with full capabilities, working in an isolated context.

  • reviewer

    Reviews code for quality, security and maintainability.

  • verifier

    Independently checks that delegated work is really done by running the real checks.

  • security-auditor

    Audits a project or PR diff for vulnerabilities and reports only what it verified in the code.

Command: /intake

Turns a vague request into a plain-language questionnaire.

Where Laya decides how to run a task, /intake handles tasks that are not ready to run, such as "change the calculation". It reads the task, analyzes your project's code and writes a Word questionnaire with every question the authors need to answer, in words a 12-year-old understands. It asks only what the code cannot answer, implements nothing and changes no file except the document.

writes .relay/intake/<date>-<title>.docx
Read the /intake docs
[ 02 // HOW IT WAS BORN ]

Switching models never fixed the agent

Coding agents fail in predictable ways. They forget a constraint you stated three messages ago, edit files when you only asked a question, report "done" when the tests failed, and drown in stale tool output. A better model does not remove these failures. Relay was built to handle them in the one layer that stays put: the harness.

01

A fork of pi

Relay was inspired by and built on pi, the minimal terminal coding agent by Mario Zechner and Earendil Works. Its spirit stays: a small core, strong defaults, and everything else as extensions.

02

The harness core

The four pillars land in the agent loop. They act on its hooks and events, not on a provider API, so they hold for Claude, GPT, Gemini, Grok or a local model.

03

laya/auto and /intake

A typo fix should not cost as much as a distributed-systems design. Laya routes each request to the right tier, and /intake turns a vague request into a plain-language questionnaire.

04

Laya learns from you

The router trains from your own sessions and runs in Docker, with its model published on Hugging Face. Version 1.0.3 ships all of it on npm.

[ CREDITS // UPSTREAM ]

Built on pi

Relay is a fork of pi, and most of what makes it work comes from there: the agent loop, the terminal UI, the provider layer, sessions, extensions, MCP and code mode. Relay keeps pi's MIT license and changelog and credits Mario Zechner and Earendil Works as contributors. On top of that foundation it adds the harness core, the laya/auto router, /intake and its own releases on npm.

pi on GitHub
[ 03 // OPEN SOURCE ]

Relay Harness

MIT
$0 open source / forever
Interactive, print, JSON and RPC modes
Harness core with all four pillars
laya/auto router and /intake
Extensions, skills, prompt templates and MCP
TypeScript SDK and the libraries behind it
npm install -g --ignore-scripts @relay-harness/coding-agent
Read the installation guide

Node.js 22.19+ • npm • Nix • from source

[ 04 // WHAT IT AIMS FOR ]

The goal: make the model a replaceable part

The rules that keep work correct should not change when you switch providers. Swap Claude for GPT, Gemini or a local model and the harness stays the same.

— Model independence

Spend quota where it matters. Easy work goes to fast models, risky work gets a strong tier and review, and failures escalate on their own.

— Cost-aware routing

Stay minimal. Strong defaults in the core; sub-agents, plan mode and the rest as extensions you write, install or share as packages.

— A small core
[ 05 // FAQ ]

Common questions

What do I need to install it?

Node.js 22.19 or newer. Install the package globally from npm and run relay in any folder. Nix and builds from source are also supported.

How is it different from pi?

Relay is a fork of pi and keeps its core. It adds the model-independent harness core, the laya/auto adaptive router, the /intake questionnaire and its own release line on npm.

Which models can I use?

Hosted providers through an API key, subscriptions through /login, any compatible endpoint as a custom provider, and local models through llama.cpp. Or let laya/auto choose for you.

Do I need Docker?

Only for the Laya model, which runs in a local container on 127.0.0.1:8737. Without Docker, laya/auto falls back to keyword rules and the rest of Relay works as usual.

Is Relay a sandbox?

No. It runs with your user's permissions. The harness asks before destructive or external commands, which guards against mistakes but is not a security boundary. Use a container for real isolation.