Picking a histogram bin count for store dashboards

Bold text Pick Better Bins beside an isometric histogram of gray columns split into bins by sliding rails.

An order-value histogram on a WooCommerce dashboard shows two tidy peaks when the range is the last 30 days, and one fat hump when it’s the last 12 months. It’s the same store and the same metric both times. The chart draws roughly the same number of bars in each view, so the bin count means something different from one to the other, because the number of orders underneath grew by more than a factor of ten.

That’s usually when somebody asks me which bin count is correct. Fetze Pijlman takes the question seriously in a Towards Data Science piece on choosing bins mathematically: treat every candidate resolution as a model, give the bin weights a Dirichlet prior, score each model as accuracy minus complexity, and average the candidates by their probability instead of declaring a single winner. On his sample data, 8 bins was the most likely model, but weighting across resolutions from 1 up to 1024 bins produced a smoother density with uncertainty bands you can draw. A store dashboard doesn’t need the full Bayesian treatment. What I take from it is the middle claim, that resolution follows the data, and that committing to one fixed number is a modelling decision you’re probably making by accident.

Why does the chart change shape when I change the date range?

Because the bin count is fixed and n is not. A few hundred orders over ten bins is a fair summary; a hundred thousand over the same ten bins is a low-resolution blur that can hide a bimodal order distribution entirely, and a split like that is something a store owner would want to see. The classic rules encode this. Sturges, log2(n)+1 bins, is the default in R’s hist() and the NumPy docs describe it as too conservative for larger, non-normal datasets. Freedman and Diaconis put the bin width at 2 × IQR / n^(1/3), so bins narrow as data grows. NumPy’s ‘auto’ setting takes the narrower of the two and tends to switch around a thousand observations, per its documentation. None of that reasoning survives into a typical custom dashboard, where someone typed 10 into the bucket loop and moved on.

WooCommerce’s own analytics screens don’t hit this, because they’re time series bucketed by day, week, or month, as the Analytics and Sales Reports docs describe. The trap is custom charts, the ones where you chose the bins yourself: order totals, cart values, product prices, time on site.

What should the histogram bin count be?

Compute it from the data, never hardcode it, and bound it on both ends so an odd month can’t produce three bars or three thousand. Freedman-Diaconis with a floor and a cap is my default for order data:

/**
 * Freedman-Diaconis bin edges for order totals, with a floor and a cap.
 */
function bbioon_order_value_bins( array $totals, int $floor = 4, int $cap = 40 ): array {
	sort( $totals );
	$n = count( $totals );

	$quantile = static function ( float $p ) use ( $totals, $n ): float {
		return $totals[ (int) floor( $p * ( $n - 1 ) ) ];
	};

	$iqr  = $quantile( 0.75 ) - $quantile( 0.25 );
	$span = $totals[ $n - 1 ] - $totals[0];

	if ( $span > 0 && $iqr > 0 ) {
		$width = ( 2 * $iqr ) / pow( $n, 1 / 3 );
		$bins  = (int) min( $cap, max( $floor, (int) ceil( $span / $width ) ) );
	} else {
		$bins = 1;
	}

	$edges = [];
	for ( $i = 0; $i <= $bins; $i++ ) {
		$edges[] = round( $totals[0] + ( $span * $i / $bins ), 2 );
	}

	return $edges;
}

// Cache per period so the axis stays stable between page loads.
set_transient( 'bbioon_bins_' . $period_key, $edges, HOUR_IN_SECONDS );

The cap matters because order totals have long tails, and one wholesale order can stretch the range until every retail order lands in the first bin. The floor is for coupon-heavy stores, where totals cluster tightly and the IQR can collapse toward zero. Cache the edges per date range rather than per page load, so the same period always renders the same axis and week-over-week screenshots stay comparable. The source article pushes further and argues you shouldn’t commit to any single bin count, because averaging over resolutions also gives you the variance to draw as bands. I think that’s more than a store dashboard needs, but if the histogram feeds a decision, like picking a free-shipping threshold, the chart has to ask a question rather than just count, and one confident shape can’t carry that by itself.

If you want a chart like this on a client store and the binning, caching and query parts sound like work you’d rather hand off, this is the kind of plugin I build, so send me the report you have in mind.

Where should the binning run?

Server-side, on a schedule. Pull order totals in a scheduled event or a WP-CLI command, compute edges and counts in PHP, park the result in a transient, and let the admin page render cached JSON. Binning in the browser means edges drift whenever the data does, and the expensive order query runs on every page load, which is one of the usual reasons analytics reports come out slow. Precomputed counts also keep the REST response small, and on shared hosting that matters more than which estimator you picked.

One caveat I’ll admit to: I’m not sure a histogram is the right chart at all for stores that sell a handful of fixed price points. If every order is one of four prices, a bar per price is more honest than any bin rule, and no estimator will tell you that.

So go find the hardcoded number. Run grep -rn "bin" wp-content/plugins/your-dashboard/ (or search your theme for the bucket loop), then switch the date range on your order chart. If the shape changes as much as the data does, the histogram bin count is probably a fixed number somebody typed once.

author avatar
Ahmad Wael
I'm a WordPress and WooCommerce developer with 15+ years of experience building custom e-commerce solutions and plugins. I specialize in PHP development, following WordPress coding standards to deliver clean, maintainable code. Currently, I'm exploring AI and e-commerce by building multi-agent systems and SaaS products that integrate technologies like Google Gemini API with WordPress platforms, approaching every project with a commitment to performance, security, and exceptional user experience.

Leave a Comment