All field notes
Case study/July 24, 2026/6 min read

From Sky Bits to Proof: A Mid-Market Fuel Operator's Evaluation Diary

Part of "Ground Truth: Stories From the Field" — proof from real commodity settings.

Every operations leader running a legacy dispatch system knows the quiet dread. The software still works, mostly. But it was built for a world with fewer trucks, cheaper fuel, and looser tolerances on shrinkage. You suspect there's money leaking out of your loads and your meter reconciliation, but you can't prove it — and you certainly can't justify a rip-and-replace on a hunch.

The operator in this story felt exactly that. Running 1,000+ trucks across a multi-state fuel distribution footprint, they'd outgrown a platform they'd nicknamed "Sky Bits" — a cloud tool that promised the world and delivered dashboards nobody trusted. When their team came to us, they didn't want a demo. They wanted a proof of concept on their own data, with numbers they could defend to a CFO.

Here's what they tested, what broke, and what finally convinced them.

The starting problem: three symptoms, one root cause

The operations leader laid out three recurring pains in the first call:

  • Loads that didn't optimize. Dispatchers were building routes by instinct and tribal knowledge. The legacy system's "optimization" was a suggestion box everyone ignored.
  • Meter reconciliation that never closed clean. The gap between what left the terminal and what arrived at the customer was chronic — and nobody could say how much of it was measurement error versus real loss.
  • No trustworthy baseline. Every KPI came with an asterisk. When leadership asked "are we improving?", the honest answer was "we think so."

The root cause underneath all three: the data was there, but it wasn't reconciled. Load plans, meter readings, and delivery confirmations lived in separate places and never met in a single ledger. That's the steady line missing under everything they moved.

Why they insisted on a POC (and why we agreed)

Buyers replacing a legacy system have been burned before. They've sat through polished demos on vendor sandbox data and then watched the promises evaporate in production. This team wanted the opposite: a controlled evaluation on their trucks, their terminals, their messy historical records.

We agreed on three conditions before writing a line of code:

  1. Real data only. No synthetic loads, no cherry-picked routes. They handed over [add time range] of historical dispatch and meter data.
  2. A pre-agreed scorecard. We defined success metrics before the POC started, so nobody could move the goalposts later.
  3. An honest reporting rule. If something didn't work, it went in the report. This wasn't a sales exercise dressed as a test.

That third condition mattered more than the first two. It's why the results held up when they went to their board.

What they tested

The POC ran across two workstreams over [add duration], scoped to a representative slice of the fleet.

Workstream 1: AI load optimization

The question wasn't "can AI build a route?" Any tool can draw lines on a map. The question was: does the optimizer respect real-world constraints and still find savings a human dispatcher missed?

They stress-tested it against:

  • Compartment and product constraints — multi-product loads with strict separation rules
  • Terminal lifting windows and rack availability
  • Driver hours and domicile realities
  • Customer delivery windows that don't bend for algorithms

The bar to clear: recommendations their most experienced dispatchers would actually accept, not override.

Workstream 2: Meter reconciliation

This was the one they cared about most, because it's where the money hides.

We ingested terminal meter readings, load tickets, and delivery confirmations and ran them through a single reconciliation ledger. The goal was to separate measurement noise from genuine variance — to tell them, load by load, where the gap was a rounding artifact and where it was a real discrepancy worth investigating.

What broke first (and why that built trust)

The first reconciliation run flagged more anomalies than the operator expected — and several of them turned out to be our mapping errors, not their losses. A handful of terminals used a legacy unit convention our ingestion didn't catch on the first pass.

We put that in the report. Verbatim.

The operations leader later told us that moment was the turning point. A vendor willing to document its own miss on day one is a vendor whose successes you can believe. Proof you can trust is proof that includes the ugly parts.

Once the mapping was corrected, the ledger closed cleaner than anything their legacy system had ever produced — and the residual variance that remained was, for the first time, explainable.

What convinced them

By the end of the evaluation, the scorecard told a clear story. The specifics belong to the operator, but the shape of the wins was consistent:

  • Load optimization surfaced route consolidations their dispatchers agreed were valid — reducing empty miles by [add %] on the tested lanes.
  • Reconciliation isolated real variance from measurement noise, giving them a defensible shrinkage figure for the first time. The suspected leakage was [add finding].
  • A single baseline replaced the asterisked KPIs. When leadership asks "are we improving?", the answer now has a number behind it.

None of these came from a slide. They came from the operator's own data, measured against metrics the operator set.

The blueprint: how to run your own evaluation

If you're weighing a switch off a legacy platform, here's the framework this operator used — steal it.

  1. Write the scorecard before the POC. Define what "good" means in numbers, and get both sides to sign it. This single step kills most vendor-buyer disputes before they start.
  2. Insist on your own data. Sandbox demos prove nothing about your terminals, your product mix, or your historical mess. Hand over a real, representative slice.
  3. Scope narrow, measure deep. A representative fraction of your fleet beats a shallow test across everything. Depth builds the confidence a board needs.
  4. Make honest reporting a contract term. Ask the vendor to document what didn't work. How they handle their own misses tells you everything about the partnership.
  5. Separate noise from signal in reconciliation. Don't accept a single "variance" number. Demand that the tool tell you which gaps are measurement artifacts and which are real.
  6. Judge optimization by dispatcher acceptance. A route your best dispatcher would override is a route that doesn't count.

The takeaway

Replacing the system that runs your loads is never a leap of faith you take lightly. This operator didn't. They demanded proof on their own ground, wrote down what success looked like, and held the vendor to it — including the parts that went wrong.

That's the standard we think every fuel logistics leader should hold. Not louder promises. Not prettier dashboards. Just a steady, reconciled line under everything you move — proven on your data, told honestly.


Want the full evaluation scorecard template this operator used? [Reach out] and we'll walk you through running the same POC on your own fleet.

See it run your workflow

Bring a real problem from your operation. Leave with a plan.

Book a demo