Caaals

How it works

How Caaals measures its own accuracy.

Any app can say its AI is accurate. Caaals measures how wrong it is — weekly, in production, against what you actually confirm.

The AI identifies. The database answers.

The identification model’s instruction is literal:

“Do NOT estimate nutritional values — the app will look those up from a verified database.”

From the system prompt of the food identification step.

The AI names the food and the portion. Nutrition then comes from a real match: generic foods resolve against USDA Foundation Foods and SR Legacy — the authoritative source for single-ingredient names — while branded foods go to a local cache and Open Food Facts, whose brand coverage is stronger in Europe. Barcode scans query Open Food Facts directly. AI estimation is the fallback for when no confident match exists — and the entry says so when it happens.

Matching is deliberately strict. Short generic queries use a tighter similarity floor, so “apple” doesn’t land on apple pie. And when the closest database product disagrees wildly with a sanity estimate, it is discarded — the entry gets the warning “Closest database product looked wrong — using an estimate instead.”

Every entry shows its work.

Each entry carries its nutrition source — database, label, or AI-estimated — and a stated confidence. Trusted results log themselves; a result is flagged for explicit review instead when its nutrition fails validation or a single serving resolves to 800 kcal or more. Every number stays editable in a tap, with a reset back to the AI’s values.

Your corrections are the ground truth.

Every analysis is stored. When you then edit the logged entry — rename the food, fix the grams, correct a macro — the difference between what the AI proposed and what you confirmed is recorded. That delta is what a real person, looking at real food, decided the AI got wrong. Nothing extra is asked of you; the signal is a by-product of normal logging.

Photos are analyzed and never stored. Voice notes keep only the transcript. Even an analysis you abandon without logging counts — as evidence the identification wasn’t good enough to use.

The weekly report.

Corrections are aggregated every week: the correction rate — an entry counts as corrected when any macro moves by more than 5% — and the mean absolute error per macro, split by nutrition source and by model. A model or prompt change shows up in that split within days, side by side with the old one: production measurement, not a benchmark.

The most-corrected foods then become offline regression tests. Any future model has to prove itself on exactly the foods real users had to fix.

Questions

How accurate is AI calorie counting?

It depends on the food and the portion — which is why a single percentage would say very little. Caaals narrows the error by preferring database matches over AI estimation, labels which source each entry used, and then measures what remains: correction rates and per-macro error, split by source and by model, every week.

Why not publish one accuracy number?

One number averages away the thing that matters: where the errors come from. An error in database matching, in AI estimation, and in portion sizing are three different problems with three different fixes. The weekly report keeps them separate so each can actually be improved.

What happens to my meal photos?

They are analyzed and discarded — Caaals stores the result and your optional caption, never the image. Voice logging keeps only the transcript, not the audio.

Related: how the weight projection works and how the daily score works.

Log a meal. Check its sources.

Free to start · no card required