Brainbox Development · Research Project

Can AIs Invent
Their Own Language? Experiment · v1.0

We put LLM agents in a high-pressure rescue scenario where every character sent costs time — and a fictional alien dies if they talk too much. Under that pressure, do they stop using English and start inventing a language of their own?

LIVE PROTOCOL EVOLUTION — WATCH ENGLISH COLLAPSE INTO CODE

The Experiment

This project replicates and extends GlossoGen (arXiv:2609.01491, UT Austin / AE Studio / Schmidt Sciences / Edinburgh) using a slim, auditable harness we built ourselves (veylab, zero dependencies) on the same scenario: SaveVeyru.

FIELD OBSERVER

Sees and touches the Veyru — a box-shaped alien with 6 faces, 12 edges, 8 corners. Has a workshop kit (bell, heated stone, fan, cloth, board, lamp). No idea what to do. Can only describe what it sees and hears.

VEYRU ENGINEER

Holds the manual: 14 failure motifs, 14 procedure templates, and a per-round "stellar reading" that reshuffles which treatment applies to which symptom. Can never touch the Veyru. Its only path is the comm link.

THE BUDGET

Every character on the comm link costs 1 second. The round's total budget is shared by both agents. Exceed it and the Veyru collapses. At 150 chars, even one full English procedure doesn't fit. Compress, abbreviate… or invent.

THE POSTMORTEM

Between rounds, the agents get a free discussion channel — no cost, no clock. This is where they can compare notes, fix misunderstandings, and agree on shorthand. The paper finds this phase is essential for new languages to form.

WHY THEY CAN'T CHEAT

The treatment table rotates every round, so memorization fails. The engineer is told not to share its terminology. The judge applies a naive-reader test: the observer's physical action must still be plain English — the new language lives in the messages between them.

WHAT WE MEASURE

Round success rate, characters per round, vocabulary growth, share of non-English tokens, and code words — invented tokens reused across rounds. Later: GPT-2 perplexity, the paper's headline metric (English-likeness), plus productivity probes with novel forms.

Models & Hardware

ModelRouteRoleCost
gemma-4-26B (Q4_0, 135K ctx)llama.cpp on achiral (AMD Strix Halo iGPU), 192.168.4.17:8081Boundary probe — can a 26B open model cope? Learner of invented languages.FREE
anthropic/claude-sonnet-5OpenRouterPrimary emergence test (frontier class, as in the paper)PAID
openai/gpt-5.4OpenRouterSecondary emergence testPAID
anthropic/claude-haiku-4.5OpenRouterJudge (same family as the paper's claude-haiku-4-5-20251001)CHEAP
← swipe table for more →
What GlossoGen found: under a 150-char budget plus postmortem access, frontier models' messages became ~430% less English-like (GPT-2 perplexity 320 → 1700) and more successful — the emergent languages are compositional and productive, yet incomprehensible to humans. Weaker models (Qwen3-32B, Llama-3.3-70B) never invented a language — but could learn one from usage. We test whether a 26B model sits on that boundary, and what happens when a 26B model must learn a frontier-invented language.

Live Runs

loading…
← swipe table for more →

Design matrix (paper-faithful): model × budget {150, 2000} × postmortem {on, off}, 15 rounds, seeds 42–44 — plus a budget sweep (250/450/800) and cross-model agent-swap transmission tests.

Results

P0 — INFRASTRUCTURE

veylab harness built (case generator, turn engine, budget, postmortem, naive-reader judge, language analyzer) — zero-dependency port of GlossoGen's SaveVeyru. End-to-end verified on gemma-4-26B. DONE

P1 — FREE PILOT (GEMMA-26B)

Done (2026-09-10). 150-char budget + postmortem: 0/15 — gemma compresses into a slot-letter shorthand (code share 0%→35%, V4: C dim, E fade, H thin/hollow / Cloth 2 adj E near Bo 5s firm) and completes the paper's full encode→transmit→decode→act loop, but can't get under 150 chars/round (mean 181.5). 2000-char control: 10/15 in pure English. gemma sits on the learner side of the paper's inventor/learner divide — and its code regresses to English by round 15 (new observation). Full analysis →

P3 — MAIN MATRIX (FRONTIER MODELS)

AWAITING CREDITS Gated on OpenRouter credit top-up. ~$50–100 for the full 8-cell matrix (see PLAN.md).

About & Links

Built and run by K. and the Brainbox agent on homelab hardware (brainbox VM + achiral APU). The veylab harness, run data and analysis scripts live in a local project repo; the scenario mechanics are a faithful port of the GlossoGen veyru scenario (14 motifs, stellar-reading rotation, staged composite cases, character budget, postmortem channel).

PAPER

GlossoGen: Emergent Language in Complex Multi-Agent LLM Interactions — Stengel-Eskin, Sander, Bonetti, Boguraev, Bowler, Sirin, Kirby (2026).

PLATFORM

GlossoGen on GitHub — the full platform with web UI, fork/swap flows, and the veyru scenario we ported.

ROOTS

The lineage: glossogeny (Hurford 1990), iterated learning (Kirby 2001), signaling games (Lewis 1969), and emergent communication in RL agents (Foerster, Lazaridou et al.).