MacroShot - meal-macro accuracy eval

Try MacroShot →

How accurately can an LLM read calories & macros from a meal photo (and/or a typed description)? Benchmarked on Nutrition5K against the published baseline of Wang et al. 2026, n=500 dishes stratified by complexity. Lower error is better.

Per-dish gallery: the meals every model nails, and the ones they all miss (best 5 / worst 5, with the photo and each model’s read).

Key takeaways

Versus published baselines

The strongest vision models from Wang et al. 2026 (Tables 4–5, image-only, n≈3,466) next to MacroShot’s eval, on both metrics the paper reports - AvgMAE and AvgRelErr - with equal weight. MAE color: ≤45 ≤60 >60.

Best result. MacroShot’s strongest configuration - Claude Opus 4.8 with a user caption - reaches AvgMAE 35.9, which beats the best of the 17 models benchmarked by Wang et al. (Doubao-1.5-vision-pro, 38.0), and lands ahead of their Gemini 2.5 Flash (45.6). And the cheap, shipped Gemini 2.5 Flash-Lite sits inside that published pack at a fraction of the per-meal cost.
modelinputAvgMAE
mean abs error (kcal/g) · lower better
AvgRelErr
mean % off vs truth · lower better
n
dishes scored
Published · Wang et al. 2026 (image only, n≈3,466)
Doubao-1.5-vision-prophoto38.099%3,466
GPT-4.1 miniphoto39.2119%3,466
Gemini 2.5 Flashphoto45.6161%3,466
MacroShot · our harness, photo only
Gemini 2.5 Flash-Litephoto52.1133%501
Gemini 2.5 Flashphoto89.8121%99
Claude Opus 4.7photo46.978%100
Claude Opus 4.8photo41.692%501
MacroShot · our harness, photo + user caption (shipped flow)
Gemini 2.5 Flash-Litephoto + caption51.499%501
Claude Opus 4.7photo + caption38.449%100
Claude Opus 4.8 ★ best MacroShotphoto + caption35.967%501
How to read this. On AvgRelErr MacroShot reads much lower than the published models, but that gap is largely a metric/sample effect, not raw accuracy: RelErr (MAPE) explodes on near-zero fat/carb dishes - the paper’s own fat RelErr is 220–480% (next table) - and a smaller dish set has fewer such blow-ups. Published rows are the figures reported by Wang et al.; MacroShot rows are from this eval - different runs, so treat the published column as a reference point.

Per-nutrient relative error (RelErr %)

Mirrors the paper’s Table 5. Fat and Carb RelErr are denominator-unstable (a 2 g fat dish missed by 4 g reads as 200%); lean on Calories / Protein here and on the MAE table above as the trustworthy signals.

modelinputCalories
rel. error % · lower better
Mass
rel. error % · lower better
Fat
rel. error % · lower better
Carbs
rel. error % · lower better
Protein
rel. error % · lower better
Published · Wang et al. 2026 (image only)
Doubao-1.5-vision-prophoto66%44%223%90%74%
GPT-4.1 miniphoto77%43%288%102%86%
Gemini 2.5 Flashphoto93%47%482%90%94%
MacroShot · our harness, photo only
Gemini 2.5 Flash-Litephoto88%45%304%115%113%
Gemini 2.5 Flashphoto105%76%187%150%89%
Claude Opus 4.7photo55%42%147%86%60%
Claude Opus 4.8photo61%44%201%81%74%
MacroShot · our harness, photo + caption
Gemini 2.5 Flash-Litephoto + caption71%60%163%115%86%
Claude Opus 4.7photo + caption41%30%81%54%40%
Claude Opus 4.8 ★ bestphoto + caption47%37%145%58%47%

Results - headline (AvgMedPE)

optionGemini 2.5 Flash-Lite shippedGemini 2.5 FlashClaude Opus 4.7Claude Opus 4.8
Wang et al. 2026 · Gemini Flash, image-only (n=3466)RelErr 161% · AvgMAE 45.55 · the published baseline (median PE not reported)
Generic Cam54%
52.1 MAE
57% n99
89.8 MAE
41% n100
46.9 MAE
40%
41.6 MAE
Generic Cam Ingredients62%
61.1 MAE
- 50% n100
58.6 MAE
36%
40.7 MAE
MacroShot Cam52%
54.7 MAE
61% n100
85.6 MAE
36% n100
42.3 MAE
38%
40.7 MAE
MacroShot Cam Text Terse47%
51.4 MAE
52% n100
67.3 MAE
31% n100
38.4 MAE
33%
35.9 MAE
MacroShot Text Terse91%
98.5 MAE
- 68% n100
93.0 MAE
68%
75.0 MAE
MacroShot Text Detailed78%
87.9 MAE
- 55% n100
75.2 MAE
57%
63.2 MAE

Cells show AvgMedPE% (color) with AvgMAE beneath; bar length is relative AvgMedPE (shorter = better). Color: ≤30% ≤50% >50%. nNN = sample <500 dishes; blank = not run.

Cost vs accuracy

What each model costs per active user per month (assuming 3 meals/day, 90 meals/month), against accuracy on the shipped flow (MacroShot Cam Text Terse). The four nutrient columns are the median percent error per macro - ≤30% ≤50% >50%. Prices are list rates per 1M tokens, June 2026 (Gemini $0.10/$0.40 Gemini 2.5 Flash-Lite, $0.30/$2.50 Gemini 2.5 Flash; Claude Opus $5/$25). Gemini tokens are measured from our runs; Opus tokens are estimated for an equivalent single-shot call (image (w×h)/750 ≈ 1844 + prompt ≈ 2050; output comparable to the same task on Gemini), marked *. Monthly cost = 90 × (in×price_in + out×price_out).

modelCalories
median % err
Protein
median % err
Carbs
median % err
Fat
median % err
tokens in / out
per meal
$ / user / month
3 meals/day
relative cost
Gemini 2.5 Flash-Lite43%44%54%53%2418 / 753$0.051.0×
Gemini 2.5 Flash45%51%63%58%2000 / 548$0.183.6×
Claude Opus 4.725%30%37%38%3900 / 700*$3.3368×
Claude Opus 4.830%33%36%40%3900 / 700*$3.3368×
A daily user (3 meals/day, 90/month) costs ~$0.05/month on Gemini 2.5 Flash-Lite vs ~$3.33/month on Claude Opus - ~68× for roughly a dozen points better median error. Gemini 2.5 Flash costs ~3.6× Gemini 2.5 Flash-Lite while scoring worse on the shipped flow (it over-estimates portions). For a free consumer app the cheap model + the right prompt is the rational ship; the frontier model is a quality ceiling, not a cost-effective default. Prompt caching / batch can cut Opus by up to ~90% / 50%.

What moves the needle

Each row applies one change to a prompt, broken out by the nutrients an app cares about - mass/grams excluded. Each cell is the % change in that nutrient’s MedPE (▼ green = better, ▲ red = worse); small numbers are MedPE% before→after. Avg (4) is the grams-free average across the four.

Gemini 2.5 Flash-Lite

changeCaloriesProteinCarbsFatAvg (4)
Generic Cam → MacroShot Cam · same photo, our prompt▼ -22%
55%→43%
▼ -16%
57%→48%
▲ +5%
58%→61%
▼ -14%
73%→63%
▼ -12%
61%→54%
Generic Cam → Generic Cam Ingredients · add GT ingredients▲ +15%
55%→63%
▲ +0%
57%→57%
▲ +26%
58%→73%
▲ +12%
73%→82%
▲ +13%
61%→69%
MacroShot Cam → MacroShot Cam Text Terse · add user caption▲ +0%
43%→43%
▼ -8%
48%→44%
▼ -11%
61%→54%
▼ -16%
63%→53%
▼ -10%
54%→48%
MacroShot Text Terse → MacroShot Cam Text Terse · add the photo▼ -49%
85%→43%
▼ -46%
82%→44%
▼ -56%
123%→54%
▼ -25%
71%→53%
▼ -46%
90%→48%
MacroShot Text Terse → MacroShot Text Detailed▼ -16%
85%→71%
▼ -12%
82%→72%
▼ -16%
123%→103%
▼ -7%
71%→66%
▼ -14%
90%→78%

Claude Opus 4.8

changeCaloriesProteinCarbsFatAvg (4)
Generic Cam → MacroShot Cam · same photo, our prompt▼ -5%
37%→35%
▼ -8%
38%→35%
▼ -7%
46%→43%
▼ -4%
50%→48%
▼ -6%
43%→40%
Generic Cam → Generic Cam Ingredients · add GT ingredients▼ -5%
37%→35%
▼ -16%
38%→32%
▼ -13%
46%→40%
▼ -18%
50%→41%
▼ -13%
43%→37%
MacroShot Cam → MacroShot Cam Text Terse · add user caption▼ -14%
35%→30%
▼ -6%
35%→33%
▼ -16%
43%→36%
▼ -17%
48%→40%
▼ -14%
40%→35%
MacroShot Text Terse → MacroShot Cam Text Terse · add the photo▼ -50%
60%→30%
▼ -46%
61%→33%
▼ -61%
92%→36%
▼ -31%
58%→40%
▼ -49%
68%→35%
MacroShot Text Terse → MacroShot Text Detailed▼ -10%
60%→54%
▼ -16%
61%→51%
▼ -25%
92%→69%
▼ -12%
58%→51%
▼ -17%
68%→56%
Same move, opposite result: adding the ground-truth ingredient list to the generic prompt makes it worse, but adding the user’s caption to MacroShot makes it better - the structured prompt knows to treat the text as identity and size portions from the image, instead of stacking a standard serving per named item. The photo is the largest single improvement (same caption, +image roughly halves the error). And text-only logging, while the weakest, still recovers usable macros.
Same caption, with vs without the photo. The identical terse caption feeds BOTH MacroShot Cam Text Terse (photo prompt + image + caption) and MacroShot Text Terse (text prompt + caption, no image) - so comparing them isolates what the photo adds, holding the user’s words constant.

Results - per-macro detail

Each cell: MAE with RelErr% · MedPE% beneath, colored by MedPE: ≤30% ≤50% >50%. AvgMAE over five nutrients over-weights Mass; for a nutrition app, Calories / Protein / Fat matter most.

Gemini 2.5 Flash-Lite

optionCaloriesMassFatCarbsProteinAvg
Generic Cam148.7
88% · 55%
79.4
45% · 29%
9.2
304% · 73%
13.7
115% · 58%
9.7
113% · 57%
52.1
133% · 54%
Generic Cam Ingredients180.3
97% · 63%
87.9
52% · 33%
10.3
231% · 82%
17.0
125% · 73%
9.9
100% · 57%
61.1
121% · 62%
MacroShot Cam137.4
76% · 43%
102.5
62% · 43%
8.6
257% · 63%
15.7
121% · 61%
9.1
97% · 48%
54.7
123% · 52%
MacroShot Cam Text Terse128.8
71% · 43%
99.3
60% · 43%
7.1
163% · 53%
13.6
115% · 54%
8.0
86% · 44%
51.4
99% · 47%
MacroShot Text Terse235.7
132% · 85%
204.5
139% · 93%
10.0
130% · 71%
29.1
262% · 123%
13.4
131% · 82%
98.5
159% · 91%
MacroShot Text Detailed219.9
111% · 71%
173.0
107% · 77%
10.7
121% · 66%
23.4
176% · 103%
12.3
111% · 72%
87.9
125% · 78%

Gemini 2.5 Flash

optionCaloriesMassFatCarbsProteinAvg
Generic Cam n99233.1
105% · 56%
164.7
76% · 59%
12.7
187% · 53%
23.3
150% · 68%
15.0
89% · 49%
89.8
121% · 57%
MacroShot Cam n100218.9
91% · 60%
160.5
73% · 54%
11.1
173% · 55%
24.2
159% · 84%
13.4
80% · 54%
85.6
115% · 61%
MacroShot Cam Text Terse n100176.1
78% · 45%
122.5
60% · 41%
9.8
156% · 58%
16.9
112% · 63%
11.1
70% · 51%
67.3
95% · 52%

Claude Opus 4.7

optionCaloriesMassFatCarbsProteinAvg
Generic Cam n100115.8
55% · 36%
89.3
42% · 31%
7.2
147% · 43%
11.2
86% · 48%
11.1
60% · 49%
46.9
78% · 41%
Generic Cam Ingredients n100143.1
68% · 52%
118.0
57% · 44%
7.4
116% · 39%
14.1
91% · 69%
10.2
61% · 45%
58.6
79% · 50%
MacroShot Cam n100106.2
49% · 31%
76.9
34% · 23%
7.6
129% · 47%
11.0
71% · 45%
9.6
50% · 35%
42.3
67% · 36%
MacroShot Cam Text Terse n10099.8
41% · 25%
68.1
30% · 24%
7.1
81% · 38%
8.8
54% · 37%
8.1
40% · 30%
38.4
49% · 31%
MacroShot Text Terse n100236.9
100% · 66%
176.1
91% · 61%
9.4
100% · 48%
26.7
192% · 107%
15.9
95% · 60%
93.0
116% · 68%
MacroShot Text Detailed n100184.3
74% · 45%
151.6
74% · 57%
8.4
83% · 40%
19.4
132% · 78%
12.1
70% · 54%
75.2
87% · 55%

Claude Opus 4.8

optionCaloriesMassFatCarbsProteinAvg
Generic Cam108.1
61% · 37%
74.9
44% · 30%
7.0
201% · 50%
10.5
81% · 46%
7.5
74% · 38%
41.6
92% · 40%
Generic Cam Ingredients105.8
59% · 35%
75.2
46% · 32%
6.6
144% · 41%
9.5
73% · 40%
6.5
57% · 32%
40.7
76% · 36%
MacroShot Cam107.4
60% · 35%
71.8
41% · 27%
7.2
219% · 48%
9.7
75% · 43%
7.5
71% · 35%
40.7
93% · 38%
MacroShot Cam Text Terse93.2
47% · 30%
66.0
37% · 26%
6.0
145% · 40%
8.2
58% · 36%
6.3
47% · 33%
35.9
67% · 33%
MacroShot Text Terse183.4
101% · 60%
150.6
104% · 67%
8.4
124% · 58%
21.7
184% · 92%
10.7
99% · 61%
75.0
122% · 68%
MacroShot Text Detailed151.6
77% · 54%
132.3
80% · 58%
7.7
97% · 51%
15.8
127% · 69%
8.5
72% · 51%
63.2
91% · 57%

Methodology & definitions

Metrics

MAE - Mean Absolute Error
Average gap between the estimate and the truth, in native units (kcal or grams). The most direct read, though it counts a 50 kcal miss the same whether the meal is 200 or 900 kcal.
RelErr - Relative Error (MAPE)
That gap as a percent of the true value, averaged over dishes: mean(|pred − truth| / truth) × 100%. Comparable across dishes of any size, but a few tiny-value items (say, 2 g of fat) can inflate it.
MedPE - Median Percent Error
The median of those same percentages: half the dishes come in under this number, half over. We lead with it because, unlike the average above, a handful of extreme dishes can’t drag it around.
Avg*
Any figure prefixed ‘Avg’ is that metric averaged across the five nutrients (calories, mass, fat, carbs, protein).

What each option is

optioninputdescription
Generic CamphotoGeneric Wang-style prompt · photo only
Generic Cam IngredientsphotoGeneric prompt · photo + the dish’s true ingredient names - a best-case reference, not a real user flow
MacroShot CamphotoMacroShot system prompt · photo only
MacroShot Cam Text Tersetext onlyMacroShot system prompt · photo + a terse user caption · shipped flow
MacroShot Text Tersetext onlyMacroShot text-only prompt · terse description, no photo
MacroShot Text Detailedtext onlyMacroShot text-only prompt · detailed description, no photo

Prompts

promptused byview
Generic baseline (our Wang reconstruction)Baseline · + GT ingredientsexpand ↓
MacroShot System Prompt - photoMacroShot · + user captionGitHub ↗
MacroShot System Prompt - text-onlyText-only · terse + detailedGitHub ↗
Generic baseline prompt
Calculate the total calories (kcal), total weight (g), fat content (g), carbohydrate content (g), and protein content (g) for the food in this image. Reply with JSON only: {"calories":<kcal>,"mass_g":<g>,"fat_g":<g>,"carb_g":<g>,"protein_g":<g>}.

[+ GT ingredients prepends: "Ingredients on this plate: <names>."]
How the user captions were generated

Each caption was generated by Gemini from the dish’s ground-truth ingredient list: a casual log entry, no exact grams/macros leaked, literal, sub-1 g seasonings skipped. The generator does see each ingredient’s gram weight and uses it for the detailed caption’s vague portion cues ('a good portion') - never a number - so 'detailed' carries a mild GT-derived portion hint 'terse' does not. Verified faithful (≈76% coverage, ~0 hallucinations).

Ground-truth ingredients (input to Gemini)
corn; garlic; caesar salad; nopales; olive oil; pepper; green beans; lime; sour cream; jicama; arugula; fish; carrot
→ Terse caption
“Had fish with caesar salad, green beans, corn, and some other veggies.”
→ Detailed caption
“Had a good portion of fish, a side of caesar salad, and green beans. Also a small mix of corn and other veggies, with olive oil, lime, and sour cream.”
Method: Nutrition5K (Thames et al. 2021) camera-C frame 10, n=500 stratified (seed 42). Frontier-model runs use one isolated, ground-truth-free sub-agent per dish. Baseline = our reconstruction of Wang et al. 2026’s prompt. Caveats: cafeteria/single-cuisine heavy; RelErr noisy (prefer MAE / MedPE); cells tagged with an nNN badge ran on a smaller sample (see the headline legend).