We should talk about RAG, because most developers still run the basic “Chunk-Embed-Match” routine and then wonder why their AI hallucinates on a 150-page credit agreement. I have spent 14 years building enterprise systems, and the lesson that keeps repeating is that structure usually matters more than the raw data. Flatten a document into 500-token chunks and you don’t just lose context, you break the document’s logic. That is where Proxy-Pointer RAG comes in.
The standard “flat” approach to retrieval-augmented generation fails because meaning in enterprise documents does not sit in isolated blocks. It lives in sections, hierarchies, and cross-references. If your system pulls a clause on “Events of Default” but misses the “Exceptions” section thirty pages later, the comparison report is worse than useless. It is a liability. Proxy-Pointer RAG fixes this by treating vector matches as “pointers” back to a structural skeleton.
Why flat RAG is a bottleneck
Most RAG setups are a game of semantic darts. You throw a query at a vector database and hope the nearest neighbors hold the answer. In a complex legal contract or a technical spec, though, the answer is usually spread across several sections. I have watched developers try to hack around this by bumping up chunk overlap or throwing a massive context window at it, and that mostly adds noise and drives up token costs.
For more on improving basic retrieval, see my earlier post on Hybrid Search and Re-ranking: Fixing RAG Accuracy. Even hybrid search falls short if the underlying architecture doesn’t know that Section 4.2 is part of Section 4.
The Proxy-Pointer RAG architecture
The idea is simple: structure-aware retrieval. Instead of searching for text alone, we search for positions inside a hierarchical tree. The architecture splits into three tiers, which keeps the semantic localization from getting lost in the noise.
- Upstream extraction: a parser like LlamaParse turns raw PDFs into clean Markdown. We then build a
_structure.jsonmap that works as the document’s GPS. - Core comparison engine: Stage 1 retrieval finds the relevant sections in the first document, then those sections act as “pointers” for a targeted cross-retrieval in the second document.
- Downstream presentation: the results don’t just land in a chat window. They go into a side-by-side comparison report written through an analytical persona, such as a senior legal counsel.
Hierarchical indexing logic
Here is a simplified look at how you would structure the hierarchical index. Instead of storing text alone, we store “breadcrumb” embeddings that carry the weight of their parent headers.
// Example of a _structure.json mapping
{
"document_id": "emerson_credit_2026",
"hierarchy": [
{
"header": "Article VII: Events of Default",
"level": 1,
"pointer": "vec_789",
"children": [
{
"header": "7.01 Termination of Commitments",
"level": 2,
"pointer": "vec_790"
}
]
}
]
}
With Gemini-3-flash handling the re-ranking, we can check whether the retrieved “pointers” match the user’s criteria before committing to a full document comparison. That staged approach keeps the pipeline both fast and accurate.
War story: the 100-page disaster
A few months back I was refactoring a document processing tool for a B2B client that compared restaurant leases. Their existing “flat” RAG kept missing the “Cure Periods” because those definitions were buried in an appendix while the default clauses sat in the first ten pages. The system would grab the default clause and ignore the grace period. That is a million-dollar hallucination waiting to happen.
Moving them to a structure-aware retrieval pipeline was more than a performance upgrade. It changed how they thought about their data. It is why I keep saying your inference architecture matters more than the LLM selection. Even the best model cannot fix a broken retrieval strategy.
Implementing it on GitHub
The full implementation is open source. Clone the Proxy-Pointer GitHub repository and you can be running in about five minutes. It uses FAISS for vector storage and is built to scale horizontally.
If this Proxy-Pointer RAG work is eating your dev hours, I can take it on. I have been wrestling with WordPress and enterprise integrations since the 4.x days.
The future of enterprise retrieval
The “vibe check” era of AI development is over. If you want to build tools that professionals actually use, whether they are lawyers, researchers, or underwriters, you cannot lean on luck. You need a system that respects the document’s hierarchy. Proxy-Pointer RAG gives you a way to turn messy source files into something you can act on. Stop flattening your docs and start pointing at them.