We put LLM agents in a high-pressure rescue scenario where every character sent costs time — and a fictional alien dies if they talk too much. Under that pressure, do they stop using English and start inventing a language of their own?
This project replicates and extends GlossoGen (arXiv:2609.01491, UT Austin / AE Studio / Schmidt Sciences / Edinburgh) using a slim, auditable harness we built ourselves (veylab, zero dependencies) on the same scenario: SaveVeyru.
Sees and touches the Veyru — a box-shaped alien with 6 faces, 12 edges, 8 corners. Has a workshop kit (bell, heated stone, fan, cloth, board, lamp). No idea what to do. Can only describe what it sees and hears.
Holds the manual: 14 failure motifs, 14 procedure templates, and a per-round "stellar reading" that reshuffles which treatment applies to which symptom. Can never touch the Veyru. Its only path is the comm link.
Every character on the comm link costs 1 second. The round's total budget is shared by both agents. Exceed it and the Veyru collapses. At 150 chars, even one full English procedure doesn't fit. Compress, abbreviate… or invent.
Between rounds, the agents get a free discussion channel — no cost, no clock. This is where they can compare notes, fix misunderstandings, and agree on shorthand. The paper finds this phase is essential for new languages to form.
The treatment table rotates every round, so memorization fails. The engineer is told not to share its terminology. The judge applies a naive-reader test: the observer's physical action must still be plain English — the new language lives in the messages between them.
Round success rate, characters per round, vocabulary growth, share of non-English tokens, and code words — invented tokens reused across rounds. Later: GPT-2 perplexity, the paper's headline metric (English-likeness), plus productivity probes with novel forms.
| Model | Route | Role | Cost |
|---|---|---|---|
gemma-4-26B (Q4_0, 135K ctx) | llama.cpp on achiral (AMD Strix Halo iGPU), 192.168.4.17:8081 | Boundary probe — can a 26B open model cope? Learner of invented languages. | FREE |
anthropic/claude-sonnet-5 | OpenRouter | Primary emergence test (frontier class, as in the paper) | PAID |
openai/gpt-5.4 | OpenRouter | Secondary emergence test | PAID |
anthropic/claude-haiku-4.5 | OpenRouter | Judge (same family as the paper's claude-haiku-4-5-20251001) | CHEAP |
What GlossoGen found: under a 150-char budget plus postmortem access, frontier models' messages became ~430% less English-like (GPT-2 perplexity 320 → 1700) and more successful — the emergent languages are compositional and productive, yet incomprehensible to humans. Weaker models (Qwen3-32B, Llama-3.3-70B) never invented a language — but could learn one from usage. We test whether a 26B model sits on that boundary, and what happens when a 26B model must learn a frontier-invented language.
Design matrix (paper-faithful): model × budget {150, 2000} × postmortem {on, off}, 15 rounds, seeds 42–44 — plus a budget sweep (250/450/800) and cross-model agent-swap transmission tests.
veylab harness built (case generator, turn engine, budget, postmortem, naive-reader judge, language analyzer) — zero-dependency port of GlossoGen's SaveVeyru. End-to-end verified on gemma-4-26B. DONE
Done (2026-09-10). 150-char budget + postmortem: 0/15 — gemma compresses into a slot-letter shorthand (code share 0%→35%, V4: C dim, E fade, H thin/hollow / Cloth 2 adj E near Bo 5s firm) and completes the paper's full encode→transmit→decode→act loop, but can't get under 150 chars/round (mean 181.5). 2000-char control: 10/15 in pure English. gemma sits on the learner side of the paper's inventor/learner divide — and its code regresses to English by round 15 (new observation). Full analysis →
AWAITING CREDITS Gated on OpenRouter credit top-up. ~$50–100 for the full 8-cell matrix (see PLAN.md).
Built and run by K. and the Brainbox agent on homelab hardware (brainbox VM + achiral APU). The veylab harness, run data and analysis scripts live in a local project repo; the scenario mechanics are a faithful port of the GlossoGen veyru scenario (14 motifs, stellar-reading rotation, staged composite cases, character budget, postmortem channel).
GlossoGen: Emergent Language in Complex Multi-Agent LLM Interactions — Stengel-Eskin, Sander, Bonetti, Boguraev, Bowler, Sirin, Kirby (2026).
GlossoGen on GitHub — the full platform with web UI, fork/swap flows, and the veyru scenario we ported.
The lineage: glossogeny (Hurford 1990), iterated learning (Kirby 2001), signaling games (Lewis 1969), and emergent communication in RL agents (Foerster, Lazaridou et al.).