Can AI write code? Test it against a benchmark

Yes, AI can write your code, and for the tedious half of my week it already does. Its accuracy tracks how much public code exists for the problem, which makes it strong where documentation and forum answers are plentiful, and quietly unreliable in corners where they aren’t. So before I trust the output, I want to know which kind of problem I’ve handed it.

Junior Jumbong wrote a piece on Towards Data Science about a study that tested this properly. The study by Winberg et al. appeared in Health Economics Review in January 2026. Instead of stopping at whether a model can produce code that runs, it checks whether the output matches known-good results, which is the part I care about.

What the study measured

The researchers had ChatGPT-4.0 Pro implement three causal inference methods from the textbook Causal Inference: The Mixtape. The methods were difference-in-differences, inverse probability treatment weighting, and regression discontinuity, in Python, R, and Stata. Reference environments were Python 3.13, R 3.6.0, and Stata 18. Four researchers worked the prompts, two with PhDs and two PhD candidates. One handled each language, and one repeated every task across all three languages with a fresh account and no conversation history.

They scored five things: accuracy against the benchmark outputs, as a binary match; efficiency, comparing command counts against the reference code; whether the code errored on execution; how much editing it needed even when it ran cleanly; and whether the same prompt produced consistent code across sessions.

The editing indicator is the one that matters most to me. Code that runs and is still wrong is the normal failure in my line of work too. It doesn’t crash and it doesn’t log anything, so nobody notices until a total is off.

Where the training data is thin

ChatGPT produced accurate code and results for most tasks in Python and R. Stata was less reliable. The authors and Jumbong both frame it cautiously, as a case study on selected problem sets rather than a universal ranking. The likely mechanism is volume. There is far more public Python and R code for a model to have learned from, and Stata code lives in places models don’t see as often.

WordPress has the same geography. Current patterns, block registration, REST controllers, wc_get_orders, are reasonably well covered. The legacy patterns are covered too well. Ask a model for order meta code and you’ll often get get_post_meta() and update_post_meta() with an order ID, or a WP_Query on shop_order. That was correct for years, and it still runs, but WooCommerce 8.2 moved orders into custom tables with HPOS, default for new stores since October 2023. Those calls write to a copy that is kept in sync rather than the real source, and sync on read is no longer something you can count on, which I wrote about when 10.7 turned it off.

Deprecated filters are the same story. woocommerce_get_price has been deprecated since WooCommerce 3.0 in favor of wc_get_price_to_display, and you will still be offered the old one, fluently and with confidence. The HPOS recipe book in the WooCommerce developer docs lists exactly which functions to stop using. The model knows the language fine. It just learned the 2016 edition of it, and your store may not be running that edition.

Two ways to accept generated code

There are two ways people accept AI-generated code, and the study puts a number on the difference between them.

The first is review by reading. You paste the output, skim it, and if it looks idiomatic you ship it. Earlier studies in this space judged code subjectively, by whether it seemed correct. That works when you already know the domain and the change is small. When you don’t, your review is an impression, and these models are very good at producing code that makes a good impression.

The second is the study’s approach: compare against a benchmark. You put known data in and check that the expected numbers come out. In store work that means a staging copy, a fixed set of orders, and totals you verified before the refactor. Then you run the generated code against the same inputs and diff the results.

I’d use the first on greenfield work that only touches current APIs, where a mistake usually shows up as a visible error. I’d insist on the second on any legacy store, anything mid-migration to HPOS, and anything that touches prices, tax, or totals. Those failures don’t throw exceptions. A wrong rounding rule in a price filter can run for months and surface as a reconciliation question from an accountant.

A checksum before you refactor anything

Before I let generated code touch order data, I snapshot a few numbers I can compare afterwards:

<?php
function bbioon_order_checksum( $status = 'completed' ) {
    $orders = wc_get_orders( array(
        'status' => $status,
        'limit'  => -1,
    ) );

    $total = 0.0;
    foreach ( $orders as $order ) {
        $total += (float) $order->get_total();
    }

    return array(
        'status' => $status,
        'count'  => count( $orders ),
        'total'  => round( $total, 2 ),
    );
}

Run it before the change and after. If the count or the total moves, the generated code changed behavior, whatever it looks like on the screen. It works on both storage backends because it goes through the CRUD layer, which is also what I want the generated code to do.

The second check is a plain grep for legacy calls before pasting anything in: get_post_meta or update_post_meta near an order ID, WP_Query with shop_order, raw SQL against wp_posts or wp_postmeta. Finding them means the code was written for a store older than yours. Once your plugin is clean, declare HPOS compatibility for it so WooCommerce stops flagging it.

If you want help setting up this kind of guardrail on a store and checking generated code against real order data before it ships, that’s work I take on, and I enjoy it.

Where I let it write, and where I don’t

These days I hand it WP-CLI scripts, one-off data fixes wrapped in checksums, test scaffolding, and rewrites between current APIs. It writes good docblocks and it never gets bored of tedious migrations. I don’t hand it anything that decides how money is calculated, or anything running on a stack I can’t benchmark first. That boundary has more to do with the benchmark existing than with how smart the model is.

I’ve written before about spec-driven development on existing WordPress sites, and this is the same idea from the testing side. Having a spec or a checksum to compare against is what makes it safe to move faster.

The question I still can’t answer is whether the thin corners shrink or just move. As HPOS-era code accumulates, the models will get better at it, but every major change, block checkout, new storage, new REST routes, resets a slice of what they know back to confident and outdated. So the useful test for a store owner is whether the person shipping it can tell when it’s writing for a site from three years ago.

author avatar
Ahmad Wael
I'm a WordPress and WooCommerce developer with 15+ years of experience building custom e-commerce solutions and plugins. I specialize in PHP development, following WordPress coding standards to deliver clean, maintainable code. Currently, I'm exploring AI and e-commerce by building multi-agent systems and SaaS products that integrate technologies like Google Gemini API with WordPress platforms, approaching every project with a commitment to performance, security, and exceptional user experience.

Leave a Comment