Meal composition: what separates good food from good macros
Working research document, copied from the author’s notes on 2026-09-06. Rough, long, and unedited apart from removing personal details. The finding page summarizes it.
Meal composition quality: what actually separates good food from good macros
Research file for the ~800-meal prepared-meal corpus.
Written 2026-09-05. Builds on [path removed]
and the v2 scorer at [path removed].
Citation discipline. Claims sourced from the literature are marked [V] when a
subagent fetched the source in this session, [R] when recalled and not re-verified.
Claims marked [M] were measured in this session against our own 816-meal corpus —
those are the strongest claims in this document because they are about our data.
Nothing here is a fabricated citation; where evidence is absent it says so.
0. The headline, stated plainly
The core question was: two meals, identical macros, one is grilled chicken + broccoli
- quinoa, the other is a small piece of chicken in cream sauce over white rice — what measurable signal separates them?
We answered this empirically rather than theoretically, by finding real pairs in the corpus whose entire nutrition panel matches and then measuring what is still different.
[M] Matched-pair experiment. Take the 434 meals with a complete panel (kcal, protein, sodium, fiber, saturated fat, sugar) and a real serving weight. Normalise all six axes per calorie, then find every pair of differently-named meals whose six axes all agree within a tolerance. Then measure the residual gap on the axes the panel does not contain:
| Panel-match tolerance | pairs | median residual gap in energy density | in potassium/kcal | in Na:K |
|---|---|---|---|---|
| 15% | 12 | 4% | (n=3, too few) | (n=3) |
| 20% | 60 | 11% | 24% | 35% |
| 25% | 156 | 10% | 27% | 33% |
When the panel matches, potassium still disagrees by 2.5×–3× more than energy density does. Potassium is the single measurable that most strongly distinguishes two meals that the nutrition panel says are the same meal.
That is not a coincidence. Potassium is the one number on a nutrition panel that tracks intact plant and unprocessed animal tissue mass and is essentially impossible to counterfeit with sauce, oil, refined starch or added protein. Cream, butter, white rice, refined flour and sugar are all potassium-poor. Broccoli, quinoa, beans, potato, squash and leafy greens are potassium-dense. Fiber captures part of this, but only the plant part and only the fermentable/structural part — potassium also picks up fish, meat and dairy tissue, and it picks up potatoes and squash that carry little fiber.
A worked pair from the corpus, panel-matched within 10% on all six axes:
[meal-service comparison removed; it is a separate private evaluation]
470 kcal | 32 g P | 870 mg Na | 7 g fibre | 4.0 g SF | 6 g sugar | 395 g | ED 1.19 | K = 820 mg
[meal-service comparison removed; it is a separate private evaluation]
510 kcal | 32 g P | 1020 mg Na | 8 g fibre | 4.5 g SF | 7 g sugar | 471 g | ED 1.08 | K = 1370 mg
Every scored axis calls these equivalent. Potassium says one delivers 67% more.
The blocker is coverage, not formula. [M] We have potassium for 209 of 816 meals (26%), spread across three of the nine sources. [meal-service comparison removed; it is a separate private evaluation] The highest-value action in this whole document is a data-acquisition task, not a scoring change: harvest potassium for the remaining 607 meals.
1. The negative results — things we tested and should NOT build
These are the most immediately useful findings because they save implementation effort. Each was measured on our own corpus.
1.1 The “added fat/sugar in the top 3 ingredients” flag adds nothing. Drop it.
This was the specific question posed. The answer is no, and the test is decisive.
[M] Restricting to the 331 weight-ordered meals with parsable ingredient lists (two of the nine sources are not weight-ordered and were excluded; [meal-service comparison removed; it is a separate private evaluation]), the flag reproduces and in fact strengthens the previously reported association:
| axis | flag = 1 (n=38) | flag = 0 (n=293) | ratio | Cohen’s d |
|---|---|---|---|---|
| saturated fat g/100 kcal | 1.90 | 0.86 | 2.20 | +1.49 |
| energy density kcal/g | 3.04 | 1.68 | 1.82 | +0.43 |
| sodium mg/kcal | 0.89 | 1.63 | 0.55 | −0.35 |
| fibre g/100 kcal | 1.70 | 1.92 | 0.89 | −0.19 |
So the flag is real. But “real” is not “useful”. The scoring question is whether it carries information the panel does not already have. It does not:
[M] OLS, energy density as outcome, n = 264 weight-ordered meals
ed ~ satfat + sugar + fibre + protein R² = 0.401
ed ~ satfat + sugar + fibre + protein + FLAG R² = 0.401
FLAG coefficient = +0.006 (se 0.131, t = +0.05)
Adding the flag to a model that already knows the panel changes R² in the fourth decimal place and its coefficient is indistinguishable from zero. The flag’s entire nutritionally meaningful content is saturated fat and sugar, which we already measure directly and better. Its residual variance is disclosure noise.
The graded version (rank position of the first added fat/sugar rather than a top-3
binary) is slightly better as a saturated-fat predictor on its own — satfat ~ fs_rank,
β = −0.031 per position, t = −2.94, but R² = 0.026. It explains 2.6% of saturated fat
variance. Saturated fat explains 100% of saturated fat. Use the panel.
Verdict: do not score ingredient position for added fat or sugar. Weight 0.
1.2 Collagen/gelatin protein inflation does not occur in this corpus. Do not build the detector.
This was flagged as “a real detection target.” [M] It is not — at least not here. Scanning all 804 meals with ingredient text, using the deepest disclosure available for each:
| protein-quality marker | meals | % of 804 |
|---|---|---|
| collagen / gelatin / hydrolysed collagen / bovine hide | 0 | 0.0% |
| any hydrolysed protein | 3 | 0.4% |
| milk protein concentrate / whey / casein | 1 | 0.1% |
| soy protein isolate / concentrate / TVP | 5 | 0.6% |
| wheat gluten / seitan | 15 | 1.9% |
| pea / rice / potato protein | 25 | 3.1% |
Zero hits on the primary target. All protein-quality markers combined touch roughly 5% of the corpus, and the largest group (pea/rice protein) is confined to the vegan meals, where it is the point of the meal rather than a cost-cutting adulterant. Those meals are also not protein outliers in a way that suggests gaming — [M] pea/rice-protein meals run lower protein density (4.12 vs 5.87 g/100 kcal) and much higher fibre density (2.61 vs 1.10) than the rest, which is the opposite of the label-inflation signature.
This makes sense structurally: label-protein inflation is a supplement and protein-bar phenomenon, driven by a nitrogen-based assay on a product whose whole value proposition is the protein number. A chicken-and-rice tray has no incentive to do it and no easy vehicle for it. The mechanism (see §4) is real; the exposure in prepared meals is not.
Verdict: the mechanism is worth knowing, the detector is not worth building. Weight 0. Revisit only if the corpus is extended to bars, shakes or “protein-boosted” SKUs — note the corpus already contains protein-shake and “double protein” items, which are the category where this would show up.
1.3 Vegetable counting from ingredient lists does not work.
[M] Across 800 meals, counting distinct vegetable tokens in the ingredient list and regressing fibre density on it:
| model | n | R² |
|---|---|---|
| fibre/100 kcal ~ distinct vegetable count | 800 | 0.007 |
| fibre/100 kcal ~ rank-weighted vegetable score (Σ 1/position) | 800 | 0.006 |
| fibre/100 kcal ~ vegetable share of ingredient list | 800 | 0.013 |
| fibre/100 kcal ~ veg count + ingredient count | 800 | 0.014 |
All coefficients are statistically significant (n = 800 will do that) and all are practically worthless. The best of them explains 1.4% of the variance in the one nutrient that vegetables are supposed to deliver. Median distinct-vegetable count is 3, p90 is 5, max is 8 — the measure barely varies.
The reason is the fundamental defect of ordinal ingredient data: an ingredient list tells you a vegetable is present, not how much. “Broccoli” as the 9th of 19 ingredients could be 5 g of garnish or 80 g of side. Position bounds it only relative to its neighbours, and after 3–4 positions the bound is vacuous.
Verdict: do not score vegetable count or variety. Fibre density and (once harvested) potassium density measure the thing you actually wanted, quantitatively.
1.4 Energy density has the best evidence and the worst discrimination.
This is the most genuinely counterintuitive finding in the document.
The v2 scorer already weights energy density at 16 (the adult) / 22 (general population),
correctly, on the strength of the NIH 4-arm factorial and the Robinson UPF meta-analysis
already cited in score.py. The causal evidence is the strongest in the whole area.
But [M] measured across our corpus, energy density is the least discriminating of every signal we have:
| signal | n | p25 | median | p75 | p75/p25 | IQR / median |
|---|---|---|---|---|---|---|
| saturated fat g/100 kcal | 503 | 0.57 | 1.00 | 1.67 | 2.94 | 1.10 |
| sugar g/100 kcal | 744 | 0.91 | 1.43 | 2.51 | 2.76 | 1.12 |
| fibre g/100 kcal | 816 | 0.70 | 1.11 | 1.73 | 2.46 | 0.93 |
| Na:K ratio | 209 | 0.63 | 0.96 | 1.52 | 2.43 | 0.93 |
| potassium mg/kcal | 209 | 1.19 | 1.67 | 2.56 | 2.15 | 0.82 |
| protein g/100 kcal | 812 | 4.21 | 5.80 | 7.50 | 1.78 | 0.57 |
| sodium mg/kcal | 814 | 1.08 | 1.50 | 1.91 | 1.77 | 0.55 |
| energy density kcal/g | 434 | 1.16 | 1.35 | 1.60 | 1.38 | 0.32 |
And against the published Rolls bands, [M] 85% of the corpus falls in a single region:
< 1.25 kcal/g ("low" ) 150 34.6%
1.25 – 1.75 221 50.9%
1.75 – 2.25 52 12.0%
> 2.25 kcal/g ("high") 11 2.5%
Only 2.5% of prepared meals reach the literature’s “high energy density” threshold. The published cutoffs were derived across the whole food supply — including crackers, cheese, nuts and oils at 5–9 kcal/g — and a plated, sauced, portioned tray simply cannot get there. Applying literature cutoffs to this corpus throws the signal away.
Two consequences:
- Band energy density on within-corpus quantiles, not literature absolutes. The v2
scorer’s band (
frac(ed, 1.1, 2.0)) is already much better than the Rolls cutoffs and clips only 27% of meals — the lowest clipping rate of any term. Keep it; do not “correct” it toward the published thresholds. - Do not over-trust the weight. [M] Energy density correlates +0.700 with total
meal calories in this corpus, because serving grams vary far less (CV 0.23) than
calories do (CV 0.36). Roughly half of what the energy-density term measures is
“this is a bigger tray,” which the
meal_sizeterm also measures.
Honest statement: energy density deserves its weight on causal grounds. But it will not move rankings much, and the expectation that it is the key discriminator among prepared meals is not supported by our data.
1.5 Energy density cannot be imputed for the meals that lack serving weight.
[M] 372 of 816 meals have no usable serving weight. The obvious fix is to predict energy density from the panel. It does not work well enough:
ed ~ satfat + protein + fibre n = 434 R² = 0.381 RMSE = 0.32 kcal/g
ed ~ satfat + protein + fibre + sugar n = 434 R² = 0.395 RMSE = 0.31 kcal/g
ed ~ above + vegetable count n = 426 R² = 0.416 RMSE = 0.31 kcal/g
RMSE 0.31 kcal/g against an interquartile range of 0.44 kcal/g. The prediction error is 70% of the entire middle-half spread of the thing being predicted. An imputed energy density would assign meals to the wrong quartile most of the time.
Verdict: never impute energy density. Leave it missing and let the scorer renormalise
across present components — which score.py already does correctly. The real fix is to
harvest serving weights, same as potassium.