Yes, install the Agent Toolkit for AWS, but install it for the guardrails rather than the demo. The demo is the fun part. Thomas Reid got a Codex agent to stand up an entire pipeline in just over 30 minutes, from a private Aurora database through Glue ingestion to an Iceberg table in S3 Tables, validated with Athena, and he wrote the whole run up on Towards Data Science. Whether I’d aim any of this at a client account comes down to something quieter. Per the toolkit documentation, AWS marks agent traffic with IAM condition keys and records every API call in CloudTrail, and that’s what makes an agent in an AWS account survivable.
The toolkit is an open-source AWS project with four parts. Skills are focused instruction packs for specific jobs, like creating an S3 Tables lakehouse table or debugging Lambda timeouts. Plugins bundle related skills under names like aws-core, aws-agents and aws-data-analytics. A rules file sets defaults, such as preferring infrastructure as code and checking AWS documentation when unsure. The AWS MCP Server ties it together with live documentation search, AWS APIs and sandboxed script execution through one endpoint. It works with Claude Code, Cursor, Codex, Kiro and anything else that speaks MCP.
Does the Agent Toolkit for AWS fix stale knowledge?
Mostly, and the way it does it is boring, which I mean as a compliment. Every model has a training cutoff. The article notes that GPT 5.5 arrived at the end of April 2026 with a knowledge cutoff of December 1, 2025, so that’s five months of new services, renamed APIs and rewritten docs the model never saw. On AWS that gap hurts. Reid’s example is a good one. A generic agent writes Athena DDL with a LOCATION clause, because that’s the standard pattern for external tables, but S3 Tables manages its own storage, so the correct pattern is clean SQL with the S3 Tables catalog passed through the query context. The generated statement looks reasonable. It just assumes the wrong model of how the service works.
The MCP server lets the agent search current AWS documentation instead of reciting training data, and the rules files push it to verify before it creates anything. WordPress is an easier target here, since core deprecates loudly and keeps old APIs working for a long time. AWS replaces services outright, and the old patterns still read as plausible years later. By the article’s count the platform has over 15,000 API calls, which is more than any model’s memory can keep current. The repo with the skills and setup instructions is aws/agent-toolkit-for-aws on GitHub, and installing the skills is one command: npx skills add aws/agent-toolkit-for-aws/skills.
How do you stop an agent from deleting the wrong thing?
Two condition context keys are attached to every request that passes through the managed AWS MCP Server: aws:ViaAWSMCPService, which is true for any AWS-managed MCP server, and aws:CalledViaAWSMCP, which carries the service principal, such as aws-mcp.amazonaws.com. You can use them in IAM policy to allow or deny whatever the agent reaches for. Blocking destructive S3 actions from MCP traffic looks like this:
{
"Effect": "Deny",
"Action": ["s3:DeleteBucket", "s3:DeleteObject"],
"Resource": "*",
"Condition": {
"StringEquals": {
"aws:CalledViaAWSMCP": "aws-mcp.amazonaws.com"
}
}
}
These keys only tag traffic that goes through the MCP server, which limits them. An agent working in a terminal can still run the plain aws CLI under your normal credentials, where no condition key appears and the deny never fires. The stronger control is the second one the article describes: a dedicated role with narrow permissions, a named profile created with aws configure, and AWS_PROFILE exported before you start the agent. I’d do both. The policy catches MCP traffic, and the profile bounds everything else, raw CLI calls included.
Can you see what the agent did afterwards?
Yes, through the services you’d use anyway. The MCP server publishes metrics to CloudWatch in the AWS-MCP namespace: Invocation, Success, UserError, SystemError and Throttle. A run of UserError entries usually means IAM denials or bad parameters, which is what you want to catch early rather than after the agent has repeated the mistake across a stack.
CloudTrail records each API call with the principal, the action, the time, the source IP and whether it succeeded. When a resource appears that nobody remembers creating, or a bill moves without an obvious cause, you can answer it from a log instead of guessing.
Is any of this relevant to WordPress work?
Most WordPress and WooCommerce builds never need CloudFormation, but the edges of a store often touch AWS: S3 for offloaded media or order exports, a Lambda job for a nightly sync, Bedrock if you’re wiring AI features into a plugin instead of calling model APIs directly. I’ve written about using Bedrock for WordPress AI and about running WordPress on serverless AWS for traffic spikes, and that’s the kind of work where an agent with stale training guesses at SDK methods that were replaced, or builds an IAM role far wider than the job needs.
There’s a client angle too. When a store owner already has an AWS account, the request tends to arrive as “the Lambda thing stopped” rather than as an architecture project. Knowing these guardrails exist means you can hand the agent a scoped profile in a sandbox instead of your main credentials, and show the client the CloudTrail log afterwards.
Where did the article’s test still need a human?
In the Athena step. Every CloudFormation resource reported CREATE_COMPLETE, the Glue job ran, six rows landed in the table, and the table still didn’t appear in the Athena console. The permissions existed at the database and table level but not at the catalog level, and Athena’s principal needs grants on the S3 Tables catalog itself. The agent found it in the end, by reading current AWS docs and granting ALL on the catalog resource, after some back and forth.
So a wall of green status icons proves very little. The habit worth copying from the article is checking what actually exists with aws cloudformation list-stack-resources and running a real query before calling the work finished.
If you want to try this today, create a scoped profile with aws configure –profile agent-lab, export AWS_PROFILE=agent-lab, let an agent run only in a sandbox account, and then read the CloudTrail events it generated. And if you have WordPress or WooCommerce work that needs AWS around it, and you want it scoped and audited rather than wide open, that’s work I take on, so send over what you have.