as_schedule_single_action( time() + 5, 'bbioon_agent_step', array( 'run_id' => $run_id ) );
This line is the closest thing WordPress has to an agent loop. It’s a queue entry that fires a hook five seconds from now, with a payload attached. No framework or reasoning engine sits behind it, and nothing stays alive between requests. I’ve come to think it’s the right starting point for AI agent architecture on this stack, and the rest is deciding where state lives and who is allowed to change it.
The prompt for this was a piece on Towards Data Science where Benjamin Nweke argues that most AI agents fail in production because they’re built backwards: teams start from what the agent should do, hand the model a pile of tools, and assume reasoning will glue the system together. His fix is components first, behavior second, a thin decision layer and plain-code orchestration. I agree with him, except on the runtime. His orchestration picture assumes a long-lived Python process. WordPress runs shared-nothing PHP, where the process dies at the end of every request, so the “just code” layer has to be hooks, database rows and a queue. It’s the same gap that shows up in what the beginner AI agent tutorial leaves out, where the demo skips the parts that break.
Where an AI agent architecture lives in WordPress
Mapped onto WordPress pieces, the layers he names look like this:
- The decision layer is one function. Prepared context goes in, a decision about the next step comes out. It doesn’t parse requests, retry calls or manage memory, because once it does, you can’t tell which part failed.
- The orchestration layer is Action Scheduler, which ships with WooCommerce. It gives you a persistent queue, retries and delayed continuation, and the project’s docs describe it processing millions of subscription payments a month, so the durability question is already answered.
- Tools are single-purpose callbacks on real hooks. A tool that calls an external API, writes a transient and updates a meta row in one function is the failure mode the article warns about, reproduced in PHP.
- State is one row per run, with statuses like pending, waiting and complete. Arrays and object cache entries don’t count, because the next step can run in a different request an hour later.
- Observability is a log table keyed by run_id that stores the exact prompt sent, the decision returned and the arguments each tool was called with. Nweke separates logging what happened from being able to tell whether it was correct, and in practice that separation is a schema decision. It’s far easier to make before the first run gets logged than after.
The entry point is usually a REST route. It checks permission_callback, inserts the run row, schedules the first action and returns a 202. The route never calls the model, so a slow provider never sits in front of your traffic. If you haven’t registered a custom route before, the adding custom endpoints page in the REST API handbook is the whole reference you need.
A skeleton that keeps the model thin
<?php
add_action( 'bbioon_agent_step', 'bbioon_agent_step' );
function bbioon_agent_step( int $run_id ): void {
$run = bbioon_get_run( $run_id );
if ( ! $run || 'complete' === $run['status'] ) {
return; // a retried action must not redo finished work
}
$decision = bbioon_decide( $run ); // the only model call
bbioon_log_step( $run_id, $decision );
if ( ! bbioon_apply( $run_id, $decision ) ) {
$run['attempts'] = ( $run['attempts'] ?? 0 ) + 1;
$run['status'] = 'waiting';
bbioon_save_run( $run );
as_schedule_single_action( time() + 60, 'bbioon_agent_step', array( 'run_id' => $run_id ) );
return;
}
$run['status'] = 'complete';
bbioon_save_run( $run );
}
bbioon_decide is the decision layer, and it should stay small enough to read on one screen: build context from the run row, send one request, return a structured decision. bbioon_apply is orchestration and tools, plain PHP you can unit test without a model key. Don’t skip the early return. Action Scheduler retries failed actions, which is what you want until a retry lands on a step that already succeeded and the customer gets the same email twice. The status check prevents that.
What still goes wrong
Model calls are slow, and Action Scheduler’s defaults are conservative on purpose. The project’s docs say a request keeps processing for up to about 30 seconds or until it has used 90% of available memory, in batches of 25, with one concurrent queue by default. An agent making six model calls per step can quietly back the queue up behind subscription renewals and order emails. Keep steps short, tune batch size and concurrency on real hosting, and split anything heavy into smaller actions. I wrote about splitting big WooCommerce background processing jobs in more detail a while back.
Stale state is the quieter one. A run can sit in the queue for an hour while a store owner changes a price, deletes a product or edits a preference. If you snapshot the catalog into the run row and trust it, the model gets blamed for a hallucination that was a stale read. Re-read what matters at the top of each step and store only run metadata. (I haven’t run this pattern on multisite, and the Action Scheduler docs say it doesn’t coordinate queues across sites, so treat it as per-site until you’ve checked.)
On any WooCommerce site you can reach, open WooCommerce > Status > Scheduled Actions, filter by failed, and open one log entry. That screen shows what built-in traceability looks like. An agent design should meet the same bar: open any run and read what it did, step by step. If you’re planning an AI feature for a store and want it built around the queue and the database rather than around the model, that’s the work I do. Send me the workflow and I’ll tell you where the state should live.