Semantic search for WooCommerce: stop at embeddings

Isometric podium with glowing blue orbs over product cubes, a dimmed machine behind; text reads Stop at Embeddings.

Theophano Mitsa built the same retrieval system four times on Towards Data Science, starting with TF-IDF and hand-picked weights and ending with a fine-tuned DistilBERT. The fine-tuned model is the part that stuck with me. It called one test critique expert-like after 30 epochs of training and novice-like after 100, with 0.685 confidence in the second answer. If you’re pricing semantic search for WooCommerce, that flip should worry you.

Where WordPress search sits

At roughly generation zero-point-five. The s parameter in WP_Query turns your term into MySQL LIKE checks on post_title and post_content, and you can inspect the generated SQL through the posts_search filter. Results get a relevance-ish sort that favors title matches, which is decent lexical matching. That’s fine for a blog, but thin for a catalog where people type “breathable summer blouse” and your product is called “Linen Shell Top 04”. (I haven’t measured the paid search plugins against this, so read it as a statement about core.)

What semantic search for WooCommerce buys you

Mainly one thing: matching ideas across different wording. In the article, “dramatic use of light” and “strong chiaroscuro effects” share almost no tokens, yet the Sentence Transformer placed them close together (the library is sentence-transformers). The store equivalent is “cotton summer shirt” finding a breathable short-sleeve blouse, and you get that from a pretrained model plus cosine similarity, no labels needed. What embeddings don’t buy is exactness. Vector similarity is fuzzy on purpose, so SKUs and model numbers tend to rank worse there, which is why the vector side belongs beside the lexical side.

What I’d still take from the TF-IDF era

The first method in the article scored documents with explicit features: TF-IDF similarity, keyword overlap, length normalization and recency decay, each bounded between 0 and 1 and summed with fixed weights. It’s boring, which is exactly why I like it, because you can print why result one beat result two. That thinking still earns its keep in store search, where the boosts are rules you can state out loud, like exact-SKU-first. A small example that pins exact title matches to the top of core search:

add_filter( 'posts_search_orderby', function ( $orderby, $query ) {
    global $wpdb;
    if ( $query->is_search() && ! is_admin() ) {
        $term = $wpdb->esc_like( $query->get( 's' ) );
        $orderby = $wpdb->prepare(
            "CASE WHEN {$wpdb->posts}.post_title LIKE %s THEN 0 ELSE 1 END, {$orderby}",
            '%' . $term . '%'
        );
    }
    return $orderby;
}, 10, 2 );

This makes the lexical layer behave. The semantic part belongs in a rerank step that scores the candidates core search already returned, which is more or less the hybrid setup I wrote about before.

If you’re weighing this for a store, core search plus an embedding rerank and two or three explicit boosts is the build I’d start with. It’s also work I take on.

Why I’d skip fine-tuning on your own labels

The article shows what happens when you try it small. The label flip between epoch 30 and epoch 100 came from a deliberately small set of critiques, and the author’s read is that the model only moved off fancy vocabulary after enough training. With millions of parameters and a tiny labeled set, most of what the model learns is style, and the confidence score keeps climbing whether or not the answer is right. Most stores I hear from have less labeled data than that toy dataset, usually none.

Where a store should stop

Stop at generation three. That means pretrained embeddings for recall, the lexical layer kept for exactness, and explicit boosts for the rules you can name. There’s no training loop and nothing to overfit. When a store owner says this SKU should rank first, the fix is a line of code instead of a retraining debate.

The question I can’t answer cheaply is how you tell a small model that’s confident because it learned something from one that’s confident because it memorized its training set. The article caught its flip only by training twice, at two epoch counts. Until someone has a boring answer to that, fine-tuning stays out of my store search proposals.

author avatar
Ahmad Wael
I'm a WordPress and WooCommerce developer with 15+ years of experience building custom e-commerce solutions and plugins. I specialize in PHP development, following WordPress coding standards to deliver clean, maintainable code. Currently, I'm exploring AI and e-commerce by building multi-agent systems and SaaS products that integrate technologies like Google Gemini API with WordPress platforms, approaching every project with a commitment to performance, security, and exceptional user experience.

Leave a Comment