The Linguist Compressor: When Rules Beat Forward Passes on Inference Cost
An arXiv paper shows CPU-only linguistic rule compressors matching recent neural methods on multiple QA benchmarks — with honest degradation at aggressive ratios.
Large language models have made prompt compression a standard tactic for cutting inference bills: score token importance with forward passes, drop the rest, and hope quality holds. A paper posted to arXiv on July 28 asks a cheaper question — can deterministic linguistic rules do the job without touching a model at compression time?
Jianfei Ma, Zhaoxin Feng, Emmanuele Chersoni, and Si Chen report that evolved rule sets, searched offline over lexical, syntactic, semantic, and discourse seeds, match recent learned compressors on short passages, multi-document reasoning, and dialogue-memory QA — while running on CPU alone at deployment.
The cost problem compression was supposed to solve
Most modern compressors still depend on language-model forward passes to rank tokens. That saves tokens downstream but spends compute upstream — a trade that grows painful at scale. The authors note that linguistics has long catalogued cues for what information matters in a text. Their hypothesis: those cues, operationalized as rules, may be enough.
They conduct offline evolutionary search to combine rule seeds into competitive compressors. At deployment, no LM forward pass is required; compression is pure CPU-side processing.
Dual-path evaluation
The team evaluates with a dual-path protocol balancing compression quality and reconstruction fidelity. Performance is strongest under light-to-moderate compression ratios and degrades as compression becomes more aggressive.
Evolutionary analysis reveals that effective rules fuse signals across linguistic levels. As compression tightens, strategies shift from token pruning toward sentence extraction.
Why labs should care
Inference economics dominate frontier deployment in 2026. A compressor that avoids LM scoring at runtime could matter most at the edge — high-volume API gateways, on-device assistants, and agent loops where every millisecond and dollar compounds.
The paper is tagged for EMNLP 2026 and runs 37 pages with six figures. It claims parity with recent advanced prompt-compression strategies under the authors' protocol, with honest degradation curves at aggressive ratios.
Limits and open questions
Rule-based compression inherits linguistics' blind spots: domain jargon, code-heavy prompts, and multilingual inputs may need seed rules the evolutionary search never saw.
For teams shipping agents that re-read long contexts every turn, the result is a design prompt: before adding another scoring model to the stack, try hiring a linguist.
### Sources
- arXiv — Every Time I Hire a Linguist, Inference Costs Go Down: On Linguistic Rules as Effective Prompt Compressors (July 28, 2026)
- arXiv — 2607.25335v2 update (July 29, 2026)