FA-86 · FemtoAgentDatasheet · Rev. 2026.10
femtoagent
A complete AI coding agent in 86 lines of bash. It runs on curl and jq. You describe a task, it writes scripts, and you approve them.
01Quick start
# needs bash, curl, jq and an OpenRouter key
$ git clone https://github.com/Berrry-Computer/femtoagent
$ cd femtoagent
$ export OPENROUTER_API_KEY="sk-or-..."
$ bash agent.sh # asks before each script
$ AUTO=1 bash agent.sh # runs scripts without asking
Source, tests and docs: github.com/Berrry-Computer/femtoagent
02How it works
› your task
│ append to history.json
▼
system prompt + history ──curl──▶ OpenRouter
(cache markers) │
▲ ▼
│ tool_calls?
│ │ │
│ yes no
│ ▼ ▼
│ run_script: show, ● reply
│ approve y/n/a, back to ›
│ bash -c, capture
└──── tool result ◀────┘
- ProtocolOpenRouter chat completions, native tool calls
- Default modelanthropic/claude-opus-5.5
- Memoryhistory.json, kept between sessions
- Prompt cachingmarks the system prompt and the last message
- Safetyasks y/n/a before every script
- Errorsdrops the failed turn, keeps history valid
- Testsoffline suite with a fake curl
- Variantagent-no-tools.sh (reads ```bash blocks)
03An experiment: three art videos
Out of curiosity, we wanted to see how an 86-line agent copes with a genuinely open-ended creative job, so we ran it next to Claude Code on the same briefs. It was one run per task, so read it as a field note, not a rigorous benchmark.
Both agents used the same model (anthropic/claude-opus-5.5 via OpenRouter) and got the same prompt, word for word. They ran side by side in tmux, each with its own API key and fully autonomous. The tasks were three open-ended art pieces: each agent had to compose music, write WebGL shaders, render a 30-second 1080×1080 video in headless Chrome and check the result.
Blind A/B qualityTie"identically good", 3 of 3 pairs
Total cost
Total wall time
femtoagentClaude Code
| Task | Agent | Time | Cost | Input tok | Cached | Output tok | Calls |
The prompts (identical for both agents)
04Watch the results
All six videos are exactly what each agent rendered from the prompt, with no human edits. In the blind viewing they looked equally good. Click a video to load the player.
05What we observed
- Overall it was close. femtoagent cost about 12% less in total, and Claude Code finished about 10% faster. The winner changed from task to task.
- Claude Code begins every request with about 33k tokens of built-in instructions and tool definitions. Most of that is read from cache, so it costs little, but it still shows up as much higher input-token counts.
- femtoagent made fewer, larger calls and wrote more output tokens per task, with a single bash tool and nothing else.
- femtoagent can't look at images, so it checked its frames with brightness statistics. In a blind viewing its videos still matched Claude Code's, which inspected frames visually.
06Method & caveats
- Cost is the change in each agent's dedicated OpenRouter key usage before and after each task. This is what was actually billed, measured the same way for both. Claude Code's self-reported cost differs from what OpenRouter charges.
- Tokens: femtoagent's come from the usage field of each API response, logged by a curl wrapper. Claude Code's come from its stream-json result (modelUsage, summed over models). "Calls" means API requests for femtoagent and turns for Claude Code.
- One run per task, with both agents running at the same time on one Apple M1, so rendering competed for CPU. Treat the per-task numbers as anecdotes, not statistics.
- Harness bug: in task 3 a script femtoagent ran read standard input and swallowed part of the scripted "exit", which caused one extra API call ($0.24, 7 s). That call is subtracted above, and the harness has been fixed.
- Claude Code ran with --dangerously-skip-permissions and femtoagent with AUTO=1. Neither was allowed installs or downloads. Each had to use ffmpeg and the puppeteer copy already installed.