Meaning Is Not in the Vector, It's in the Computation: A Paper Challenges RAG's Geometric Assumption
An Awkward Fact
Suppose you build a RAG (retrieval-augmented generation) system with a good embedding model. A user asks how to take a screenshot on an iPhone. You retrieve a document titled "How to take a screenshot on iPhone" with cosine similarity 0.92 — a perfect hit.
But the same corpus also contains a sentence like "How to capture the screen while taking photos on an iPhone." It means almost the same thing as the question, yet scores only 0.71 in cosine similarity — ranked 15th, truncated away.
You might think: just switch to a better embedding model.
Jiaqi Deng's paper tells you: no model will fix this. The problem isn't the model — it's the entire paradigm. Semantic equivalence never existed inside a single sentence's vector in the first place.
Two Worldviews
The paper cleanly decomposes the question "do two sentences express the same meaning":
The geometric worldview: every sentence has a vector, and "same meaning" is a geometric property of two vectors, computable via cosine similarity. The entire RAG, dense retrieval, and duplicate-detection industry rests on this assumption.
The computational worldview: meaning identity is not a property of a single sentence, but a relation computed when two sentences appear together in the same forward pass.
These worldviews aren't rhetorical. The paper designs two decisive discriminative experiments:
- Late fusion: encode both sentences separately, concatenate the vectors, and feed them to a linear probe. If meaning truly lived in vectors, a linear probe should read it out.
- Shuffle control: replace sentence B with an unrelated sentence while keeping the format identical, and run the joint forward pass. If the probe reads template format rather than pair relations, it should still work after shuffling.
The Numbers Speak
The paper's core results on PAWS-X (a dataset specifically built to test paraphrase adversarial cases):
| Model | Cosine | Linear probe (late fusion) | Joint forward | Shuffle | |------|------|----------------------|---------|---------| | Qwen2.5-14B | 0.614 | 0.615 | 0.950 | 0.496 | | Qwen2.5-7B-Inst | 0.571 | 0.612 | 0.961 | 0.526 | | Llama-3-8B | 0.530 | 0.514 | 0.897 | 0.542 | | GPT-2 XL | 0.521 | 0.513 | 0.756 | 0.457 |
Key findings:
1. Joint forward propagation crushes everything. Concatenate two sentences, run one forward pass, and read "same meaning or not" from the last token's hidden state: AUC jumps from ~0.55 to ~0.95. No fine-tuning, no training — just a linear probe.
2. The shuffle is the fatal blow. Swapping B for an unrelated sentence while keeping the format drops AUC to 0.50. The model reads genuine pair relations, not templates.
3. 1.5B is enough. Qwen2.5-1.5B achieves a joint-probe AUC of 0.947, nearly identical to the 32B model's 0.938. The capability saturates around 3B parameters.
4. GPT-2 XL already has it. At 0.76, it far exceeds cosine's 0.52. This computational ability isn't exclusive to large models — it's a basic capability of language models.
Why This Matters
Think about your RAG system: millions of documents, one query encoded as a vector, cosine similarity search over a vector store, top-k returned.
The implicit assumption: same meaning = high cosine similarity.
But the paper shows that for "same meaning, different wording" hard cases, cosine similarity is essentially chance. Retrieval-specialized models — BGE, E5, GTE, MiniLM — score only 0.55–0.65 AUC on PAWS-X.
You might say: use a reranker. The paper tested that too. BGE-reranker-large reaches 0.94, but MS-MARCO and Jina rerankers stay in the 0.55–0.64 range. Crucially, a reranker is itself a joint forward pass — it concatenates query and document through a cross-encoder. Rerankers work precisely because they unintentionally exploit the computational worldview.
A Deeper Finding: Cross-Model Consistency
The paper includes a striking experiment. Models from different families, scales, and independent training runs (Llama, Mistral, Qwen, Phi, GPT-Neo, OPT) run joint forward passes and their judgments agree highly — not just on labels, but even on errors.
This means language models trained by different teams, on different data, with different architectures have converged on the same computation for judging paraphrase. That's not an engineering coincidence — it's something deeper.
A distillation experiment reinforces this: a 1.5B student, trained on unlabeled sentence pairs scored by a teacher, recovers the same judgment capability. But no linear combination of the teacher's independent vectors can.
The relation is distillable; the vector is not.
Implications for Engineering Practice
1. The bi-encoder paradigm has a ceiling. Pre-encoding documents into vectors and searching by cosine has a hard failure mode on paraphrased queries. Fine-tuning helps on one dataset but can hurt on another (PAWS improves, QQP degrades) — you're just fitting a specific dataset.
2. Rerankers aren't a nice-to-have; they're a necessity. But they're slow and expensive — you can't run a cross-encoder over a million documents. Real systems use two stages: bi-encoder coarse filtering to top-100, then reranker fine ranking. That paradigm is right, but be aware the coarse stage may already have discarded the true paraphrase.
3. Meaning is an operator, not a vector. For tasks like paraphrase detection, don't try to read it out of a single vector. Run a joint forward pass or train a dedicated pair reader.
An Honest Assessment
Limitations:
1. Only PAWS-X was tested. PAWS is designed for the hardest case: high lexical overlap, different meanings. For more common semantic search scenarios, embedding models may be good enough.
2. The engineering cost of joint forward passes. Running query-plus-candidate through a model for every candidate is extremely expensive for RAG.
3. No new RAG architecture proposed. The paper identifies the problem without offering a replacement. Rerankers are an existing approximation; the paper hints better pair-reader architectures may exist.
Even with these limitations, the paper's value lies in using a clean experimental design to challenge an assumption an entire industry depends on. Next time you tune your embedding cosine-similarity threshold, ask yourself: are you measuring "semantic similarity," or just "phrasing neighborhood"?
---
Paper: Computation Over Geometry: Meaning Identity Is Computed, Not Shipped in the Embeddings
Author: Jiaqi Deng (independent researcher)
Core conclusion: Meaning identity is not a geometric property of single-sentence vectors but a relation computed in a joint forward pass over both sentences. The "geometric assumption" the RAG industry is built on is wrong in hard cases.