Reproducible · transparent · documented

Numbers, not promises

Zahlen, keine Versprechen

Every comparison below is a real, runnable scenario — the same job in Python + LangChain vs. Pipe. Full commands are in the docs and examples.

Jeder Vergleich unten ist ein reales, ausführbares Szenario — dieselbe Aufgabe in Python + LangChain vs. Pipe. Alle Befehle stehen in Doku und Beispielen.

14 LOC
RAG pipeline (Python: 26)
RAG-Pipeline (Python: 26)
1.2s
3 parallel LLM calls
3 parallele LLM-Calls
8.6 MB
Binary vs 345 MB venv
Binary vs 345 MB venv
55×
VM speedup on fib(20), 0.6–55× measured
VM-Beschleunigung bei fib(20), 0,6–55× gemessen
📐

RAG pipeline — code size

RAG-Pipeline — Code-Umfang

The same semantic-search question-answering pipeline: embed a knowledge base, find the nearest match, and ask the LLM.

Dieselbe semantische Such-&-Frage-Pipeline: Wissensbasis embedden, nächsten Treffer finden und das LLM fragen.

Python+LC
26 lines
26 LOC
Pipe
14 lines
14 LOC

1.9× less code — measured on the real files in benchmarks/python-vs-pipe. Same DeepSeek model on both sides; the Python side additionally needs an embedding provider (Ollama) because DeepSeek has no embeddings API, while Pipe embeds locally with zero setup.

1,9× weniger Code — gemessen an den realen Dateien in benchmarks/python-vs-pipe. Beide Seiten mit demselben DeepSeek-Modell; die Python-Seite braucht zusätzlich einen Embedding-Anbieter (Ollama), weil DeepSeek keine Embedding-API hat, während Pipe lokal ohne Einrichtung embeddet.

-- Pipe RAG (14 lines)
ai_provider "deepseek"

docs: [read_file "data/docs/database.txt"]
push docs (read_file "data/docs/caching.txt")
push docs (read_file "data/docs/api.txt")
push docs (read_file "data/docs/deployment.txt")

vectors: embed_batch docs
q_vec: embed "How do we rate-limit API requests?"
top: nearest q_vec vectors 3

context: ""
for idx in top
    context: context ++ (at docs idx) ++ "\n---\n"

answer: ask ("Context:\n" ++ context ++ "\nQuestion: " ++ "How do we rate-limit API requests?")
print answer
⚡

Parallel LLM calls — wall time

Parallele LLM-Calls — Laufzeit

Three independent LLM calls (deepseek-v4-flash). Sequential Python takes ~3.3s. Python + asyncio.gather() runs them in ~1.3s. Pipe's >> operator: ~1.2s total — with no interpreter/import startup at all.

Drei unabhängige LLM-Calls (deepseek-v4-flash). Sequentielles Python braucht ~3,3s. Python + asyncio.gather() läuft in ~1,3s. Der >>-Operator von Pipe: ~1,2s gesamt — ohne Interpreter-/Import-Startup.

Sequential
~3.3s
~3.3s
Python+LC asyncio
~1.3s
~1.3s
Pipe >>
~1.2s
~1.2s

2.8× faster than sequential. Both parallel versions are LLM-latency-bound, but the Python process needs ~7s of interpreter + import startup (langchain, openai, ...) before the first call — Pipe starts in milliseconds. ai_batch scales the same idea to hundreds of texts with built-in rate limiting.

2,8× schneller als sequentiell. Beide Parallel-Varianten sind LLM-Latenz-gebunden, aber der Python-Prozess braucht ~7s Interpreter- + Import-Startup (langchain, openai, ...) vor dem ersten Call — Pipe startet in Millisekunden. ai_batch skaliert dieselbe Idee auf hunderte Texte mit eingebautem Rate-Limiting.

-- 3 calls in parallel — ~1.2s total
a: "Löse 7*8+4 und antworte nur mit der Zahl."  >> ask
b: "Löse 12*12 und antworte nur mit der Zahl."  >> ask
c: "Löse 100/4 und antworte nur mit der Zahl."  >> ask

print a ++ " | " ++ b ++ " | " ++ c
💾

Deploy size — binary vs. Docker image

Deploy-Größe — Binary vs. Docker-Image

The real Python venv behind the examples above (langchain, langchain-deepseek, langchain-ollama, faiss-cpu, fastapi) measures 345 MB on disk — Pipe is one statically-linked 8.6 MB binary.

Das reale Python-venv hinter den obigen Beispielen (langchain, langchain-deepseek, langchain-ollama, faiss-cpu, fastapi) misst 345 MB auf der Platte — Pipe ist eine einzige statisch gelinkte 8,6-MB-Binary.

Python+LC
345 MB
345 MB
Pipe
8.6 MB
8.6 MB
Pipe (UPX)
~2.9 MB
~2.9 MB

40× smaller. Deploy to a server or Raspberry Pi with scp pipe binary. No venv, no pip, no Docker, no runtime dependencies.

40× kleiner. Deploy auf Server oder Raspberry Pi per scp pipe binary. Kein venv, kein pip, kein Docker, keine Laufzeit-Abhängigkeiten.

🏎️

Execution engine — VM vs tree-walker, measured

Ausführungs-Engine — VM vs. Tree-Walker, gemessen

Run with pipe -vm -q script.pipe to compile to bytecode and execute on a stack-based VM instead of the tree-walker. Speedup depends heavily on the workload — from 0.6× to 55×. Here are the measured medians (pipe -bench, pure execution time):

Ausführen mit pipe -vm -q script.pipe kompiliert zu Bytecode und führt auf einer Stack-basierten VM statt im Tree-Walker aus. Die Beschleunigung hängt stark vom Workload ab — von 0,6× bis 55×. Hier die gemessenen Mediane (pipe -bench, reine Ausführungszeit):

Tree-walkerVMSpeedup
fib(20) recursive366 ms6.7 ms55×
list sum 100004.5 ms1.5 ms3.0×
list push + sum 2000015.7 ms18.2 ms0.9×
fizzbuzz 1–1000.22 ms0.34 ms0.7×
string concat 2000085 ms150 ms0.6×

The VM wins big on function-call and recursion-heavy code (up to 55× on fib), but simple loops, list building and string concatenation are comparable or even slower — the tree-walker optimizes those well. Take a single "~7×" claim with a grain of salt: measure your own workload with pipe -bench.

Die VM gewinnt deutlich bei funktionsaufruf- und rekursionslastigem Code (bis 55× bei fib), aber einfache Schleifen, Listen-Aufbau und String-Concatenation sind vergleichbar oder sogar langsamer — der Tree-Walker optimiert diese gut. Ein pauschales „~7×" mit Vorsicht genießen: eigenen Workload messen mit pipe -bench.

⚙️

CPU microbenchmarks — Pipe vs Python, honest

CPU-Mikrobenchmarks — Pipe vs. Python, ehrlich

Identical programs, no LLM calls, wall time including interpreter startup (median of 5 runs). CPython wins these — Pipe is not a JIT compiler and is not designed to outrun Python on raw CPU. Where Pipe actually wins is LLM workflows: no 7-second langchain import startup, built-in parallel >> and ai_batch.

Identische Programme, keine LLM-Calls, Wandzeit inkl. Interpreter-Startup (Median aus 5 Läufen). CPython gewinnt diese Benchmarks — Pipe ist kein JIT-Compiler und nicht dafür gebaut, Python bei roher CPU-Leistung zu schlagen. Wo Pipe wirklich gewinnt, sind LLM-Workflows: kein 7-Sekunden-langchain-Import-Startup, eingebautes Parallel->> und ai_batch.

Pipe (VM)PythonFaster
fib(20) recursive72 ms18 msPython 4.1×
loop sum 1,000,000215 ms140 msPython 1.5×
string concat 20000206 ms30 msPython 6.8×
list push 2000094 ms21 msPython 4.4×

Note the startup share: Pipe's ~100 ms here includes binary startup + bytecode compile, Python's ~15 ms is pure interpreter startup. Both are negligible next to the ~7s Python import startup for langchain in the real LLM benchmarks above. Rerun: python3 tools/measure_perf.py.

Startup-Anteil beachten: Pipe's ~100 ms hier beinhalten Binary-Start + Bytecode-Kompilierung, Python's ~15 ms sind reiner Interpreter-Start. Beides ist vernachlässigbar neben dem ~7s-Import-Startup von langchain bei den echten LLM-Benchmarks oben. Neu messen: python3 tools/measure_perf.py.

Pipe vs. Python + LangChain

Pipe vs. Python + LangChain

Same job. Less code. Built-in safety.

Gleicher Job. Weniger Code. Eingebaute Sicherheit.

Python + LangChainPipe
RAG pipelineRAG-Pipeline26 LOC26 Zeilen14 LOC14 Zeilen
Sandbox LLM accessLLM-Zugriff sandboxenCustom middlewareCustom MiddlewareOne sandbox_profile blockEin sandbox_profile-Block
Switch AI providerKI-Provider wechselnRewrite SDK callsSDK-Calls umschreibenai_provider "deepseek"
Deploy to serverAuf Server deployenDocker + venv + pipDocker + venv + pipscp pipe binaryscp pipe binary
Parallel LLM callsParallele LLM-Callsasyncio.gather() boilerplateasyncio.gather()-Boilerplate>> operator, ai_batch
Binary size (with deps)Binary-Größe (mit Deps)345 MB venv345 MB venv8.6 MB8.6 MB

How these numbers were measured

Wie diese Zahlen gemessen wurden

  • All LOC and wall-time numbers above are measured on the real files in benchmarks/python-vs-pipe/ (Pipe: pipe/*.pipe, Python: python/*.py). Rerun with python3 tools/measure.py --run.
  • Alle LOC- und Laufzeit-Zahlen oben sind an den realen Dateien in benchmarks/python-vs-pipe/ gemessen (Pipe: pipe/*.pipe, Python: python/*.py). Neu messen mit python3 tools/measure.py --run.
  • Both sides use the same model (deepseek-v4-flash). The Python RAG example additionally needs Ollama nomic-embed-text for embeddings — DeepSeek has no embeddings API — while Pipe's embed_batch works locally with zero setup.
  • Beide Seiten nutzen dasselbe Modell (deepseek-v4-flash). Das Python-RAG-Beispiel braucht zusätzlich Ollama nomic-embed-text für Embeddings — DeepSeek hat keine Embedding-API — während embed_batch von Pipe lokal ohne Einrichtung funktioniert.
  • Wall times depend on provider latency; the Python figures include ~7s interpreter/import startup, Pipe includes none.
  • Laufzeiten hängen von der Provider-Latenz ab; die Python-Werte enthalten ~7s Interpreter-/Import-Startup, Pipe keinen.
  • Binary sizes: go build -ldflags="-s -w" -trimpath. UPX figure is make build-upx. Python size is the real du -sh of the benchmark venv.
  • Binary-Größen: go build -ldflags="-s -w" -trimpath. UPX-Wert via make build-upx. Python-Größe ist das reale du -sh des Benchmark-venv.
  • Execution-engine numbers come from pipe -bench (pure execution time, average of 5 runs, no startup). CPU microbenchmarks are identical programs in benchmarks/performance/, median of 5 runs, wall time incl. interpreter startup. Rerun with python3 tools/measure_perf.py.
  • Ausführungs-Engine-Zahlen stammen aus pipe -bench (reine Ausführungszeit, Durchschnitt aus 5 Läufen, ohne Startup). CPU-Mikrobenchmarks sind identische Programme in benchmarks/performance/, Median aus 5 Läufen, Wandzeit inkl. Interpreter-Startup. Neu messen mit python3 tools/measure_perf.py.
  • Run every example yourself in the browser playground — no install needed.
  • Jedes Beispiel selbst ausprobieren im Browser-Playground — ohne Installation.