Why recursive language models beat ReAct for long context

Thumbnail with the words Point Not Paste beside an isometric card tower where a small arm lifts one card by a pointer tag.

For about a year now, the answer to any long-context task has been “use a bigger context window” or “stuff more into the RAG pipeline.” Both burn real money in tokens and neither one fixes the reasoning. If you have ever watched a 128k context model flail around looking for one logic bug in a large codebase, you have seen the problem.

ReAct, CodeAct and most of the frameworks around them treat context as a buffer you keep appending to, so memory and reasoning end up competing for the same room. Recursive Language Models (RLMs) borrow an idea software has had for decades instead: pass by reference.

Why your current agent is slow

In a standard agentic harness, ReAct or CodeAct alike, the model has to read everything in order to know anything. Ask it to summarize 300 logs and it reads the logs, generates a thought, then writes a response. If that response has to carry the original data, the model regenerates every one of those tokens. That is slow and expensive, and regenerated data is where the hallucinations creep in.

RLMs work inside a REPL (Read-Eval-Print-Loop). The harness does not feed the model the whole prompt. It hands over a pointer to the data, and the model decides what to read, what to slice, and where to park intermediate results in variables that survive across turns.

I wrote earlier about a pragmatic agentic AI workflow for WordPress. RLMs push past that by separating what the model remembers from what fits in its context window.

How RLMs handle long context without “rot”

The environment, usually Deno or Pyodide, exposes a context variable. The model never sees what is inside it until it calls something like print(context[:500]). That changes what the agent can actually do with the data:

  • Programmatic exploration: it can run regex or pandas operations against the data right there in the REPL.
  • Variable composition: subagents hand back Python variables rather than text, so the parent agent can use a result without ever reading it into its own context.
  • Output length: the result lives in the environment as a variable, so the model’s completion token limit stops mattering.

Here is the same categorization task done the traditional way, then the way Recursive Language Models do it.

# Traditional CodeAct Approach
# The model must read all data and manually rewrite the JSON.
def process_logs(logs):
    # LLM reads 1MB of logs...
    # LLM hallucinates half-way through the 5,000 line JSON output.
    return formatted_json

# RLM Approach (The "Pass by Reference" Logic)
# The model writes code to handle the heavy lifting.
```repl
import json
# context is already in the environment
results = await llm_query("Extract errors from context") 
# results is a Python list variable, not a string in the prompt
FINAL(results) 
```

Why this matters for WordPress developers

Take auditing a WooCommerce site with 100+ active plugins. Feed that codebase into a standard LLM and context rot shows up fast: the model forgets the hook definitions it read at the start. With an RLM you tell the agent to map the directory structure recursively, keep the function definitions in a dictionary, and open the source of a specific function only when it needs to look.

Those variables still vanish when the session ends, which is why this pairs well with a persistent AI context layer. That combination is the difference between a demo and something you can leave running in production.

If this kind of agent plumbing is eating your dev hours, hand it to me. I have been wrestling with WordPress and AI integrations since the early days.

Stop regurgitating data

What makes Recursive Language Models work is that they are lazy: they avoid generating any token they do not have to. Push the heavy lifting into a sandboxed REPL, let recursive subagents pass variables instead of strings, and the agent can work across millions of tokens without ever holding them all in context.

author avatar
Ahmad Wael
I'm a WordPress and WooCommerce developer with 15+ years of experience building custom e-commerce solutions and plugins. I specialize in PHP development, following WordPress coding standards to deliver clean, maintainable code. Currently, I'm exploring AI and e-commerce by building multi-agent systems and SaaS products that integrate technologies like Google Gemini API with WordPress platforms, approaching every project with a commitment to performance, security, and exceptional user experience.