Two Ollama-powered agents collaborate on the same isolated workspace:
- Builder — implements the requested app by proposing complete file writes/deletes.
- Reviewer — independently inspects the resulting workspace and either approves it or sends concrete feedback.
- Orchestrator — applies changes, performs deterministic static checks, logs every round, and stops at a hard iteration limit.
No cloud model API is required.
v0.1 experimental scaffold
The first milestone is deliberately constrained. The agents can create/edit files only inside one experiment workspace. They do not receive shell access in v0.1. This lets us observe the Builder → Reviewer feedback loop before adding more dangerous capabilities.
- Python 3.11+
- Ollama
- Local models:
ollama pull qwen2.5-coder:3b
ollama pull qwen3:4bThese lighter defaults are intended to keep the first experiment responsive on consumer hardware. Larger models can still be selected from the CLI.
Verify Ollama is running:
ollama list
curl http://localhost:11434/api/tagsgit clone https://github.com/miflow13/RelayLab.git
cd RelayLab
python -m venv .venv
source .venv/bin/activate
python -m unittest discover -s tests -v
python relay.py \
"Build a tiny single-page notes app with add, complete, and delete actions." \
--name notes-v1RelayLab will create:
experiments/notes-v1/workspace/ # app produced by the agents
runs/<run-id>/ # full experiment transcript + metadata
task
↓
Builder
↓
structured file actions
↓
RelayLab applies changes
↓
static deterministic checks
↓
Reviewer
├─ changes_requested → Builder
└─ approved → done
The loop stops after 6 rounds by default even if the Reviewer never approves.
The Builder must return JSON:
{
"summary": "What changed",
"files": [
{
"path": "index.html",
"content": "<!doctype html>..."
}
],
"deletes": []
}The Reviewer must return JSON:
{
"status": "approved",
"feedback": "Why this satisfies the task."
}Ollama is requested to return JSON-formatted responses, and RelayLab validates the protocol before applying actions.
v0.1 intentionally does not give either model unrestricted computer access.
- Files are resolved against one experiment workspace.
- Absolute paths and
..path traversal are rejected. - Model-authored code is not executed.
- Python source is syntax-compiled without running it.
- JSON files are parsed without running generated programs.
- Git operations and pushes are unavailable to the agents.
- Every agent response, applied action, check result, and review is logged.
- Runs have a hard maximum-round limit.
This is still experimental software rather than a hardened security boundary. Keep generated work disposable until stronger process/container isolation is added.
| Role | Model |
|---|---|
| Builder | qwen2.5-coder:3b |
| Reviewer | qwen3:4b |
The previous 7B/8B defaults were intentionally reduced after early testing showed they could overwhelm a desktop when Ollama fell back heavily to CPU.
Override either from the CLI:
python relay.py "Build a pomodoro timer" \
--name pomodoro \
--builder-model qwen2.5-coder:7b \
--reviewer-model qwen3:8bRelayLab is meant to become an actual experiment, not just a chatbot demo. Each run records enough information to compare:
- rounds to approval
- reviewer rejection reasons
- deterministic check failures
- model combinations
- final artifacts
- whether a two-agent loop improves over a single-agent baseline
RelayLab/
├── relaylab/
│ ├── agents/
│ │ ├── builder.py
│ │ └── reviewer.py
│ ├── config.py
│ ├── jsonutil.py
│ ├── ollama_client.py
│ ├── orchestrator.py
│ └── workspace.py
├── prompts/
│ ├── builder.md
│ └── reviewer.md
├── experiments/
├── runs/
├── tests/
├── relay.py
└── pyproject.toml
- Local Ollama client
- Builder / Reviewer role separation
- Structured action protocol
- Workspace path isolation
- Hard iteration limit
- JSON run transcripts
- Deterministic static checks
- GitHub Actions unit tests
- Strong OS/container sandbox
- Controlled shell/tool calling
- Browser-based app verification
- Git checkpoints per accepted round
- Single-agent baseline mode
- Run comparison metrics
- Live web UI for watching both agents work
Experimental project. Add a license before distributing or incorporating third-party code.
