Skip to content
zyoung11Public

About

A truly minimal pi-coding-agent compiled to an 9 MB zero-dependency native executable.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Latest commit

 

History

294 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

╔══╗                  ╦        
║  ║ o                ║   o    
╠══╝ ╦      ════      ╠═╗ ╦ ╔╗╔
║    ║                ║ ║ ║ ║║║
╩    ╩                ╚═╝ ╩ ╝╚╝

A truly minimal pi-coding-agent compiled to an ~9 MB zero-dependency native executable.

Forked from pi-mono (v0.84.3) and compiled with scriptc (v0.0.35) into a single static Linux binary. No JS engine, no node_modules at runtime.

pi-bin pi (Node.js)
Binary 9.5 MB single ELF 289 MB node_modules + Node.js ≥ 24
Startup to --version ~3 ms ~380 ms
Startup to --list-models ~40 ms ~420 ms
TUI ready (no session) ~70 ms ~490 ms
TUI interactive (5 MB session) ~0.3 s ~540 ms (includes full transcript)
History back-fill (5 MB session) ~9 s in background —
Idle memory (TUI running) ~17 MB ~145 MB
Memory (chat round-trip) no measurable growth ~152 MB
Memory (5 MB session loaded) ~105 MB ~259 MB
Operating System Support Linux (currently) Linux, Windows, MacOS
Runtime dependencies libc, libm, libz Node.js ≥ 24 + npm packages

Each number is the median of repeated runs on Linux x86-64 (same config, same terminal size, same local model endpoint). Startup numbers are the local paths; a configured credential for a catalog provider adds about one second while the models.dev directory refresh runs.

On multi-MB sessions the transcript loads progressively: the newest screens render first and the TUI is interactive immediately, older history back-fills in the background over the following seconds, and the editor border shows ── ⠙ Loading session history… with input held until the fill finishes. The border also hosts the session status spinners (working, compaction, branch summary, retry), and a resumed session with history drops the new session banner so the first screen holds only the transcript.

Differences

Added:

  • A built-in llama.cpp provider for self-hosted llama.cpp routers (the upstream llama extension, reworked for the static build): /login llama.cpp records the server URL and an optional key, the model catalog and its metadata (thinking, context window, quantization, modalities) are discovered from the router, and /llama lists, loads, and unloads the models the router holds with a framed progress view. No model downloads.
  • Thinking levels are measured, not assumed. pi asks the server to render its own chat template for every level (/apply-template, no inference) and the ladder offers exactly the levels the template accepts: refusals are hidden, and levels that render an identical prompt are aliases of one depth, so only the name the template resolves to is shown. A level a template refuses at request time is dropped from the ladder at runtime instead of failing the turn.

Removed:

  • Extension system (TypeScript extensions, extension marketplace)
  • grep / find / ls
  • Fullscreen TUI mode (regular mode only)
  • Telemetry, update checks, first-run wizard
  • OAuth/subscription authentication (Claude Pro/Max, ChatGPT/Codex, GitHub Copilot logins are not available; API-key providers work fine via /login)
  • Automatic light/dark theme switching (dark + custom themes only)
  • Image resizing (images pass through as-is. Over 4 MB is rejected)
  • /changelog and the extension slash commands are gone

Everything else (the agent loop, tools, TUI, sessions, compaction, skills, prompt templates, and the bash/edit/read/write workflow) behaves the same.

Session loading differs by design. Instead of rendering the whole transcript before the TUI becomes usable, which grows with history size, the newest screens render first and the TUI is interactive immediately. Older history back-fills in the background over the following seconds with a loading placeholder in the editor border, so scrolling far up into old history may still be filling in during that window.

The scriptc compiler (v0.0.35) is bundled in the repo at scriptc/, so no separate install is needed. It is included unmodified and pinned for reproducible builds and debugging. No compiler changes were made, and all fixes live in this repo's TypeScript source.

Install

Grab the latest binary from Releases, then:

mv pi ~/.local/bin/

Building from source

Prerequisites:

  • Node.js ≥ 24 (scriptc requires it. Make sure node --version reports ≥ 24)
  • Linux x86-64
git clone https://github.com/zyoung11/pi-bin.git
cd pi-bin

npm install --ignore-scripts   # install npm dependencies
npm run build:native           # compile → ./pi (9.5 MB ELF)
npm run check                  # optional: biome + tsgo --noEmit

The build reads packages/coding-agent/src/cli.ts and emits the static binary to ./pi.

Models

Provider coverage (✓ works with /login, ✗ not supported yet):

Provider pi-bin
Ant Ling ✓
Amazon Bedrock ✗
Anthropic ✗
Azure OpenAI ✗
Baseten ✓
Cerebras ✓
Cloudflare AI Gateway ✗
Cloudflare Workers AI ✗
DeepSeek ✓
Fireworks ✓
GitHub Copilot ✗
Google Gemini ✗
Google Vertex ✗
Groq ✓
Hugging Face ✓
Kimi For Coding ✗
llama.cpp (self-hosted router) ✓
MiniMax (and CN) ✗
Mistral ✗
Moonshot AI (and CN) ✓
NVIDIA ✓
OpenAI ✗
OpenCode (and Go) ✓
OpenRouter ✓
Qwen Token Plan (and CN / Individual) ✓
Together ✓
Vercel AI Gateway ✗
xAI ✗
Xiaomi (and Token Plan AMS / CN / SGP) ✓
Z.AI (and Coding CN) ✓
  • ✗ providers need API adapters the static build does not implement: Anthropic-wire protocol (Anthropic, Kimi For Coding, MiniMax, Vercel AI Gateway), OpenAI Responses protocol (OpenAI, xAI, Azure OpenAI), Google/Bedrock/Mistral proprietary protocols, OAuth/subscription logins (GitHub Copilot, OpenAI Codex), or special credential chains (Amazon Bedrock, Cloudflare, RADIUS).
  • Self-hosted llama.cpp routers are first-class: /login llama.cpp adds one, its catalog and per-model metadata are discovered automatically, and /llama loads and unloads the models the router holds (no model downloads). Thinking levels come from the server's own template render, so the ladder offers exactly the levels a model accepts.
  • Other self-hosted OpenAI-compatible endpoints (vLLM, ...) work through a models.json file in the agent config directory. Those providers can set detectChatTemplateThinking: true to pick up the enable_thinking chat-template control and the accepted thinking levels automatically.
  • Select models with --model, the /model command, or Ctrl+P.

About

A truly minimal pi-coding-agent compiled to an 9 MB zero-dependency native executable.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Contributors

Languages