Dietary patterns and ultra-processed food epidemiology
Working research document, copied from the author’s notes on 2026-09-06. Rough, long, and unedited apart from removing personal details. The finding page summarizes it.
Ultra-Processed Food, Dietary Patterns, and Diet-Quality Scoring
Evidence review for the per-meal health score (~800 meal-delivery meals)
Date: 2026-09-05 Purpose: Decide (a) how much weight “how processed is it” should carry when scoring an individual prepared meal, and (b) whether an existing validated scoring system beats the custom score we built.
Evidence-handling rules used in this document
- Every numeric claim below is either (a) verified by fetching the source during this research session, or (b) explicitly tagged
[UNVERIFIED]. Nothing is asserted from memory without a tag. - Study design is labelled everywhere: RCT, cohort, mechanistic/animal, or expert-judgement validation.
- Effect sizes are given with CIs and, where obtainable, absolute risks. Relative risks alone are treated as insufficient for weighting decisions.
0. Bottom line up front
- Keep the custom score. Do not replace it with a whole-diet index. HEI-2020, AHEI-2010, DASH scores, Mediterranean scores and DII are all diet-level instruments. Three of them (AHEI, MDS, DII) are not merely unvalidated at meal level, they are arithmetically incoherent at meal level. See §2.
- The only credible per-item validated alternatives are Nutri-Score (2023), the FSA/Ofcom NPM, NRF9.3, and the FDA 2024 “healthy” meal criteria. Of these, Nutri-Score is the best-validated against hard outcomes and has a fully specified published algorithm (reproduced in §3). But it is 100g-based, ignores processing and additives entirely, and was not designed to rank composite ready meals against each other.
- Recommendation: keep the custom score as primary, and compute Nutri-Score 2023 + FSA-NPM as secondary reference columns for external validation and sanity-checking. Where our score and Nutri-Score disagree sharply on a meal, that meal is worth a manual look. This costs little and buys defensibility.
- The weighting is roughly right in shape but wrong in the tails. Our ~76 / 10 / 14 split is not defensible as it stands. The 14% on additives is the weakest-evidenced block in the model — no additive class has randomised hard-outcome evidence at realistic dietary exposure. And the processing block is measuring something that has now been directly quantified and found small.
- The decisive new evidence: the NIH 4-arm factorial trial (NCT05290064, n=36, results posted 2026-08-12, not yet peer-reviewed) decomposes the ultra-processing overeating effect. Of a 948 kcal/day total gap: energy density +662 kcal/day (p<0.0001), hyperpalatability +158 (p=0.023), and ultra-processing per se +128 kcal/day (95% CI −8.2 to 265.0, p = 0.065 — not significant). Robinson et al.’s meta-analysis of 10 UPF RCTs independently finds the intake effect vanishes once energy density is matched (SMD 0.02, −0.51 to 0.55).
- And the cohort category we actually care about is the least incriminated one. Across five large disaggregated cohorts, “ready-to-eat/heat mixed dishes” — our 800 meals — are null for mortality (HR 1.02, 0.99–1.05), null for CVD, null for multimorbidity, and only weakly positive for T2D. The UPF signal lives almost entirely in processed meat and sugary drinks: deleting just those two categories from the UPF variable collapses the CVD association from HR 1.11 to 1.00 (0.96–1.05).
- So: “how processed is it” should carry ~5-8%, not 10%, and should be re-scoped from a generic NOVA/ingredient-count marker to two specific, unambiguous signals — processed meat content and liquid sugar. The weight freed up should go to energy density and fibre. Full table in §4.7.
1. Ultra-processed food — the causal question
1.1 Hall et al. 2019 — the anchor trial
Hall KD, Ayuketah A, Brychta R, et al. “Ultra-Processed Diets Cause Excess Calorie Intake and Weight Gain: An Inpatient Randomized Controlled Trial of Ad Libitum Food Intake.” Cell Metab. 2019;30(1):67–77.e3. PMID 31105044.
Design (RCT, inpatient controlled feeding). n = 20 weight-stable adults, NIH Metabolic Clinical Research Unit, randomised crossover, two 14-day arms (ultra-processed vs unprocessed), ad libitum intake, meals matched on presented calories, energy density, macronutrients, sugar, sodium and fibre.
Results.
- +508 ± 106 kcal/day on the ultra-processed arm (p = 0.0001)
- Weight +0.9 ± 0.3 kg (UPF) vs −0.9 ± 0.3 kg (unprocessed)
- Eating rate +17 kcal/min, +7.4 g/min on UPF
- The excess was entirely carbohydrate (+280 kcal) and fat (+230 kcal); protein was −2 ± 12 kcal/day — null
- PYY rose and ghrelin fell on the unprocessed arm; leptin null; pleasantness ratings did not differ between arms
Correction to a common premise: there is NO corrigendum about the energy-density matching. Two errata exist and both were checked:
- Cell Metab 2019;30(1):226 (PMID 31269427) — a diet-code data-entry error on one participant’s meal sheets. The primary outcome was unaffected; corrected per-meal excesses are breakfast +144, lunch +248, dinner +108 kcal.
- Cell Metab 2020;32(4):690 (PMID 33027677) — record exists, content not fetched
[UNVERIFIED].
The real interpretive weakness is disclosed inside the paper itself, and it is serious:
| UPF arm | Unprocessed arm | |
|---|---|---|
| Presented energy density, including beverages | 1.024 kcal/g | 1.028 kcal/g |
| Presented energy density, non-beverage foods only | 1.957 kcal/g | 1.057 kcal/g |
| Energy density of food actually consumed | 1.36 ± 0.02 | 1.09 ± 0.02 |
The headline “matched for energy density” was achieved by adding energy-containing beverages to the unprocessed arm. The solid food on the UPF arm was ~85% more energy dense, and the paper’s own Discussion concedes this “likely contributed to the observed excess energy intake.”
Published critique. Astrup’s “NO” position in the AJCN 2022 debate: the effect “can be entirely explained by more conventional and quantifiable dietary factors, including energy density, intrinsic fiber, glycemic load, and added sugar.”
Other limitations: n = 20, 2 weeks per arm, single site, metabolic ward, no possible blinding.
1.2 Dicken et al. 2025 (UPDATE) — the guideline-matched trial
Nature Medicine 2025;31(10):3297–3308 (PMC12532614).
Design (RCT, free-living controlled feeding). n = 55 randomised / 50 ITT, adults in England, BMI 25 to <40, habitually ≥50% of energy from UPF. Two 8-week ad libitum arms with a 4-week washout, both diets formulated to meet the UK Eatwell Guide.
Results.
- Minimally-processed arm −2.06% body weight (−2.99, −1.13); ultra-processed arm −1.05% (−1.98, −0.13)
- Between-arm difference −1.01% (95% CI −1.87, −0.14), p = 0.024, Cohen’s d = −0.48 (≈0.96 kg)
- Fat mass −0.98 kg (p = 0.004)
- Adherence: 15.5% vs 91.2% of energy from UPF — the manipulation worked
Why it is the more relevant trial for us. Hall compared ultra-processed against unprocessed with no requirement that either be a healthy diet. UPDATE holds national dietary-guideline conformance constant in both arms and asks what processing adds on top. That is exactly our question: given two meals that both look nutritionally reasonable, does the more processed one carry extra risk? Answer: yes, but about one percentage point of body weight over eight weeks.
But it does not escape the confound either. The UPF arm was 1.60 vs 1.25 kcal/g — still ~28% more energy dense. Meeting a nutrient-based guideline does not equalise energy density. Both arms also lost weight, so this is not a real-world contrast.
1.3 The mechanism — and the factorial trial that decomposes it
This is the most important evidence in the entire review for our weighting decision.
The NIH 4-arm factorial trial (NCT05290064)
Status: completed 2025-08-15; results posted to ClinicalTrials.gov 2026-08-12. NO peer-reviewed publication was found. Treat as registry-posted results, not peer-reviewed.
Randomised crossover, n = 36 analysed (14 F / 22 M, age 36.3 ± 11.3, BMI 29.5 ± 7.0), four 1-week ad libitum inpatient diets:
| Arm | Processing | Energy density | Hyperpalatable content | Intake (kcal/day, LS mean ± SE) |
|---|---|---|---|---|
| UPF HH | ultra-processed | High | High | 3424.3 ± 108.5 |
| UPF HL | ultra-processed | High | Low | 3265.9 ± 108.5 |
| UPF LL | ultra-processed | Low | Low | 2604.4 ± 108.5 |
| UNF LL | unprocessed | Low | Low | 2476.0 ± 108.5 |
The decomposition:
| Contrast | Isolates | Δ kcal/day | 95% CI | p |
|---|---|---|---|---|
| UPF HH vs UPF HL | Hyperpalatability | +158.4 | 21.9 to 294.9 | 0.023 |
| UPF HL vs UPF LL | Energy density | +661.6 | 524.9 to 798.3 | <0.0001 |
| UPF LL vs UNF LL | Ultra-processing per se | +128.4 | −8.2 to 265.0 | 0.065 — not significant |
| UPF HH vs UNF LL | The full “UPF effect” | +948.4 | 811.9 to 1084.8 | <0.0001 |
Of the total 948 kcal/day gap: ~70% is energy density, ~17% is hyperpalatable nutrient combinations, and ~14% is everything else about ultra-processing — which does not reach significance.
Secondary outcomes: eating rate did not differ across arms (39.21 / 39.02 / 39.04 / 38.43 g/min, all pairwise p = 0.99) and palatability ratings did not differ (VAS 65.2 / 62.4 / 66.7 / 64.2, all p ≥ 0.6).
Cautions: 1-week arms, n = 36, unpublished, and the residual CI is wide — this is “not demonstrated,” not “demonstrated to be zero.”
Mechanism-by-mechanism verdict
| Mechanism | Direct human experimental support | Verdict |
|---|---|---|
| Energy density | Rolls 2006: −575 kcal/day for a 25% ED reduction (n=24) with no change in hunger ratings. NIH 4-arm: +662 kcal/day. Robinson meta-analysis: SMD 0.71 (0.45–0.98) when UPF arms were >10% higher ED, vs SMD 0.02 (−0.51, 0.55) when ED was matched | STRONGEST. This is the dominant mechanism |
| Eating rate / oral processing | Forde 2026 (AJCN), n=41, 2-week crossover: −369 kcal/day comparing UPF-Slow vs UPF-Fast at matched energy density and palatability. Robinson 2014 meta SMD 0.45 | STRONG when texture is deliberately manipulated — but did not differ across the NIH 4-arm menus. Sufficient, not always operative |
| Hyperpalatability (Fazzino-style fat+sugar / fat+salt combos) | NIH 4-arm: +158 kcal/day (22–295), p=0.023 | MODEST but real |
| Palatability as liking/pleasure | Hall 2019 pleasantness p=0.13; NIH 4-arm all p≥0.6; Robinson meta SMD 0.38 NS | NULL — people do not overeat UPF because they enjoy it more |
| Protein leverage (Simpson & Raubenheimer) | Gosby 2011 RCT, n=22: +12% intake at 10% vs 15% protein, nothing above 15%. Hall’s own ceiling estimate ~50% | REAL BUT SMALL. Population data show leverage on intake but consistently not on BMI. In Hall 2019 the protein excess was −2 ± 12 kcal/day — null |
| Additives — emulsifiers | The human CMC trial is FRESH, n=16, 11 days, 15 g/day — GI and microbiome endpoints, no intake or adiposity outcome. Chassaing’s causal chain is mouse at 1% w/v | NO experimental basis for assigning any share. Large mouse-to-human dose gap. (Note: no trial named “CADEC” was found) |
| Additives — non-nutritive sweeteners | Suez 2022 (n=120 RCT): person-specific glycaemic effects. But four RCT meta-analyses show NNS substitution reduces weight (−0.79 to −1.06 kg) | Evidence points the opposite way from the intuition |
| Food matrix | Karl 2017 (n=81): ~92 kcal/day net for whole vs refined grain. Novotny 2012: Atwater overestimates almond energy by 32% | REAL but small, and acts largely through energy density and eating rate — double-counting risk |
| Appetite hormones | Hall 2019 measured fasting values only, after 2 weeks of divergent energy balance | Consequence at least as plausibly as cause |
Does processing add risk beyond nutrient composition?
Two direct tests both say: not detectably.
- NIH 4-arm, UPF-LL vs UNF-LL: +128.4 kcal/day, 95% CI −8.2 to 265.0, p = 0.065. A UPF diet engineered to be low in energy density and low in hyperpalatable content was statistically indistinguishable from an unprocessed diet.
- Robinson et al. meta-analysis of 10 UPF RCTs (medRxiv preprint, not peer-reviewed): energy intake SMD 0.71 (0.45–0.98) when UPF arms were >10% higher in energy density, vs SMD 0.02 (−0.51, 0.55) when energy density was matched. Their conclusion verbatim: “there is no convincing evidence that ultra-processing of food has an effect on energy intake or body weight independent to nutritional profile.”
Counterweight: Dicken 2025 found ~1% more weight loss on the minimally-processed arm with both arms meeting the Eatwell Guide — but that arm was still 28% lower in energy density, so it does not escape the confound either.
Synthesis. “Ultra-processed” is a highly reliable marker for a bundle of physical food properties — high energy density, soft/fast-eating texture, and fat+sugar / fat+salt nutrient combinations — each of which causes overeating in controlled human experiments. It is not currently demonstrated to carry causal risk beyond that bundle.
Direct implication for our score: if we already measure energy density, and we already reward protein and fibre, we have captured most of the causal pathway. A separate “processing” weight is largely double-counting the same mechanism.
1.4 How much of the UPF–disease association survives adjustment for nutrient profile?
The honest answer: it depends on the outcome, and the split is sharp.
Mortality — the association does NOT survive. Fang Z, Rossato SL, Hang D, Khandpur N, et al., BMJ 2024;385:e078476 (PMC11077436). NHS (74,563 women) + HPFS (39,501 men), 34-year mean follow-up, 48,193 deaths.
| Model | Q4 vs Q1 all-cause mortality HR |
|---|---|
| Age, sex, calories only | 1.22 (1.18–1.25) |
| Full multivariable | 1.04 (1.01–1.07) |
| + AHEI-2010 (diet quality) | “attenuated toward null” |
In the 16-cell joint analysis, AHEI predicted mortality within every UPF stratum, but UPF did not predict mortality within AHEI strata. The authors’ own conclusion, verbatim: “our data together suggest that dietary quality has a predominant influence on long term health, whereas the additional effect of food processing is likely to be limited.”
CVD — the association does NOT survive, and there are two independent knockdowns. Mendoza K, Smith-Warner SA, … Mattei J, Lancet Reg Health Am 2024;37:100859 (PMC11403639). NHS + NHSII + HPFS, 16,800 CVD cases / 5,387,896 person-years.
| Analysis | CVD Q5 vs Q1 | CHD | Stroke |
|---|---|---|---|
| Age + energy adjusted | 1.21 (1.16–1.27) | 1.27 (1.20–1.35) | 1.13 (1.05–1.22) |
| Fully adjusted | 1.11 (1.06–1.16) | 1.16 (1.09–1.24) | 1.04 (0.96–1.12) |
| + modified AHEI | 1.03 (0.98–1.09) | 1.07 (1.00–1.14) | — |
| Removing SSBs + processed meats from the UPF variable | 1.00 (0.96–1.05) | 1.06 (1.00–1.13) | 0.92 (0.85–0.99) |
Read that last row carefully. Deleting two food categories from the definition of “ultra-processed food” eliminates the entire cardiovascular association. That is the single most decision-relevant result in the UPF epidemiology for our project.
Type 2 diabetes — the association DOES survive. This is the genuine exception and it should be conceded.
- EPIC (Dicken SJ et al., Lancet Reg Health Eur 2024, PMC11551512): n=311,892, 14,236 incident T2D. HR per 10% g/day from UPF: 1.17 (1.14–1.19) → 1.19 adding saturated fat/sugar/sodium → 1.14 adding a Mediterranean score → 1.15 with both. Attenuation ≈ 12% of the excess risk. Waist-to-height ratio mediated 46.4% (38.3–54.4).
- US cohorts (Chen Z, Khandpur N, et al., Diabetes Care 2023;46:1335-44, PMC10300524): 19,503 cases. Q5 vs Q1 HR 1.56 (1.47–1.65), or 1.28 (1.21–1.36) with baseline BMI. Nutrients (fibre, refined starch, added sugar, sodium, minerals, partially hydrogenated oils) collectively mediated only 11.9% (4.6–27.7%) — and only fibre and sodium were individually significant.
Also relevant: substitution analyses in EPIC — replacing UPF with minimally processed food + culinary ingredients gave HR 0.86 (0.84–0.88); replacing with processed food HR 0.82 (0.79–0.85).
No Mendelian randomisation study of NOVA-defined UPF intake on disease outcomes exists. There is no established genetic instrument for “UPF intake”. The strongest quasi-experimental tool available elsewhere in nutritional epidemiology is simply unavailable here — a genuine hole in the causal case.
Note a conflict in the literature we should represent fairly. Dicken & Batterham’s 2021 Nutrients review concluded that UPF associations “remain significant and unchanged in magnitude after adjustment for diet quality or pattern,” and Lane 2024 cites this approvingly. But that review predates both Fang 2024 and Mendoza 2024, which found substantial attenuation for mortality and CVD. The 2021 conclusion is out of date for those two outcomes; it still holds for T2D.
1.5 Lane 2024 BMJ umbrella review — what it actually establishes
Lane MM, Gamage E, Du S, et al. Ultra-processed food exposure and adverse health outcomes: umbrella review of epidemiological meta-analyses. BMJ 2024;384:e077310 (PMC10899807).
- Scope: 45 pooled analyses from 14 meta-analysis studies, n = 9,888,373. 32 of 45 outcomes showed direct associations.
- Credibility: only 4 of 45 graded “convincing” (class I) — CVD mortality RR 1.50 (1.37–1.63), T2D dose-response RR 1.12 (1.11–1.13), anxiety OR 1.48, common mental disorders OR 1.53.
- GRADE quality: ZERO high. 4 moderate, 22 low, 19 very low. 13 of 45 graded “no evidence.” Excess significance bias detected in 9 of 28 testable analyses (32%).
- No randomised-trial meta-analyses existed to include. The entire umbrella review is built on observational data.
What it establishes: that UPF intake is consistently associated with a wide range of adverse outcomes across a very large body of observational data. What it does not establish: causation, independence from nutrient profile, or that the associations are of high certainty. Note the internal tension: the one “convincing” finding that is also GRADE-moderate is T2D — the same outcome that survives nutrient adjustment. The CVD-mortality finding graded “convincing” on class criteria is simultaneously graded GRADE very low.
1.6 Critiques of NOVA as a construct
Measured reproducibility is poor — this is the strongest empirical critique.
- Braesco et al. 2022 (Eur J Clin Nutr): Fleiss κ = 0.32–0.34 among 159–177 trained food and nutrition specialists classifying foods with ingredient lists available. That is “fair” agreement at best.
- Harvard coders applying NOVA to FFQ items agreed on only 70.2% at first pass (Khandpur 2021, PMC8453454).
- PREDIMED-Plus: the same FFQ data yielded a UPF share of 7.9% under NOVA versus 45.9% under the IARC system; cross-system ICC 0.51.
If the same food data can produce a 7.9% or a 45.9% UPF share depending on which classification you apply, the exposure variable is not measuring a stable property of food.
The conceptual critique. The formal point/counterpoint is in Am J Clin Nutr 2022;116(6): “YES” (pp. 1476–81), “NO” (pp. 1482–88), consensus (pp. 1489–91) — not Nature Food. Astrup’s “NO” position is the sharpest: the Hall 2019 effect “can be entirely explained by more conventional and quantifiable dietary factors, including energy density, intrinsic fiber, glycemic load, and added sugar.” Gibney’s critique is that NOVA is defined by formulation and provenance rather than by any measurable physical or chemical property, so it cannot be operationalised objectively.
Note also: Sadler et al.’s work is qualitative/conceptual (thematic analysis of professionals’ perceptions, n=27), not a quantitative reliability study — the reliability number to cite is Braesco 2022.
The defence, represented fairly. Monteiro and colleagues argue that NOVA captures an industrial-formulation gestalt that reductionist nutrient analysis misses, that the consistency of associations across dozens of cohorts and populations is itself evidence, and that demanding nutrient-adjusted independence is over-adjustment because the nutrients are themselves consequences of the processing.
Verdict: NOVA is a useful public-health heuristic with poor edges and poor measured reliability. It is not sound enough to carry weight as a scoring input for individual products.
1.7 Do UPF subgroups differ? (Decision-critical — yes, enormously)
Five large disaggregated cohort analyses. They agree on the top of the list and disagree in the middle.
| Subgroup | Mortality (Fang, NHS+HPFS) | CVD (Mendoza) | T2D (Chen, US) | T2D (Dicken, EPIC) | Multimorbidity (Cordova, EPIC) | Consensus |
|---|---|---|---|---|---|---|
| Processed meat / animal-based | 1.13 (1.10–1.16) | harm, all 3 outcomes | harm | 2.25 (1.96–2.57) | 1.09 (1.05–1.12) | HARM — unanimous |
| Sugar- & artificially-sweetened beverages | 1.09 (1.07–1.12) | harm, all 3 | harm | 1.25 (1.22–1.28) | 1.09 (1.06–1.12) | HARM — unanimous |
| UPF breads / cereals | 1.04 (1.02–1.07) | inverse (stroke, CVD) | inverse | 0.65 (0.57–0.73) | 0.97 (0.94–1.00) | NULL to PROTECTIVE |
| Packaged sweet snacks / desserts | 0.99 (0.97–1.02); 0.94 CVD death, 0.95 cancer death | — | inverse | 0.89 (0.84–0.95) | null | NULL to PROTECTIVE |
| Yogurt / dairy desserts | 1.07 (1.04–1.10) | inverse | inverse | — | null | CONFLICTING |
| Savoury snacks | 1.01 null | inverse | inverse | 2.77 (1.09–7.05) unstable | null | CONFLICTING / mostly null |
| Ready meals / mixed dishes | 1.02 (0.99–1.05) NULL | not significant | harm (magnitude n/a) | 1.16 (1.01–1.35) | not significant | NULL for mortality, CVD, multimorbidity; weak-positive for T2D only |
This is the finding that most directly changes what we should build. The category our 800 meals belong to — ready-to-eat/heat mixed dishes — is the least incriminated of all the obviously-ultra-processed categories. It is null for mortality, null for CVD, null for multimorbidity, and only weakly positive for T2D at the boundary of significance on an extrapolated exposure scale.
A high-protein ready meal does not carry the risk the UPF label implies. Applying a flat “it’s ultra-processed” penalty to it is applying a population average that demonstrably does not describe its own subgroup.
1.8 Absolute risk — the magnitudes are small
- Mortality (Fang 2024): Q1 (median 3.0 UPF servings/day) 1,472 deaths per 100,000 person-years; Q4 (median 7.4 servings/day) 1,536 per 100,000 PY. Absolute difference 64 per 100,000 PY, i.e. 0.064 percentage points per year. NNH ≈ 1,563 people for one extra death per year, across a 2.5-fold difference in UPF intake, in a cohort averaging ~50 years old at baseline.
- CVD, youngest available cohort (NHSII, women aged 25–42 at entry, mean 36.7): 72 CVD cases per 100,000 PY — about 4.3× lower than the pooled older cohorts. Q5 vs Q1 HR 1.22 → roughly 0.44 percentage points of absolute 30-year CVD risk for the entire extreme-quintile contrast, before diet-quality adjustment (which cuts it further).
- T2D is the exception with a meaningful absolute signal: EPIC 419 cases per 100,000 PY; HR 1.17 per 10% of grams from UPF → ~71 excess cases per 100,000 PY. US cohorts ~95 excess per 100,000 PY.
Framing for a adult: relative risks of 1.04–1.30 applied to a baseline risk an order of magnitude below these middle-aged cohorts produce very small absolute effects. The one outcome worth taking seriously is type 2 diabetes.
2. Validated diet-quality scores — should we use one instead?
2.1 The single most important structural fact
Almost every well-validated diet-quality index was built, calibrated and validated on whole-diet exposure over long periods, usually from an FFQ or repeated 24h recalls. Applying them to a single meal is not merely “unvalidated” — for several of them it is mathematically undefined. There are three distinct failure modes, and it matters which one applies:
| Failure mode | What breaks | Which indices |
|---|---|---|
| Population-relative construction | The score is defined by the cohort’s median or by global reference means/SDs. A single meal has no such referent. | Mediterranean MDS (Trichopoulou), DII |
| Components that cannot exist in one meal | Alcohol, long-term trans-fat intake, “days per week of X”, total daily variety | AHEI-2010, several DASH variants |
| Density-computable but semantically wrong | Arithmetic works (per 1,000 kcal), but the score punishes a meal for not being an entire balanced day | HEI-2020, DASH (Fung-style) |
The third is the seductive one. HEI-2020 will return a number for a single meal because since 2005 its standards are density-based (per 1,000 kcal). NCI states the density design means “the HEI can be applied to the diets of individuals and to various levels of the food stream,” and lists four such levels: national food supply, food processing, community food environment, and individual food intake — note that “a single meal” is not one of the four listed levels. NCI also cautions that “an individual’s HEI score based on a given day’s intake would not necessarily reflect the score based on their usual or habitual intake” — a caution that applies a fortiori to one meal out of ~21 per week.
The practical consequence: a perfectly good high-protein chicken-and-broccoli meal scores 0/10 on Dairy, 0/10 on Whole Grains, 0/5 on Total Fruits and 0/5 on Whole Fruits, and therefore caps out around 60/100 on HEI — not because it is a bad meal but because it is not a whole day. Any index with “adequacy” components spanning all food groups will systematically punish focused, single-purpose meals. This is the core reason a whole-diet index is the wrong tool for our job.
Source: https://epi.grants.cancer.gov/hei/uses.html (fetched)
2.2 HEI-2020
What it is. USDA/NCI index measuring conformance to the Dietary Guidelines for Americans. 13 components, total 100 points, density-based per 1,000 kcal, with 9 “adequacy” components (higher = better) and 4 “moderation” components (lower = better).
Full scoring standards (fetched from https://epi.grants.cancer.gov/hei/hei-2020-table1.html):
| Component | Max pts | Standard for max score | Standard for zero |
|---|---|---|---|
| Total Fruits | 5 | ≥0.8 cup-eq/1,000 kcal | none |
| Whole Fruits | 5 | ≥0.4 cup-eq/1,000 kcal | none |
| Total Vegetables | 5 | ≥1.1 cup-eq/1,000 kcal | none |
| Greens & Beans | 5 | ≥0.2 cup-eq/1,000 kcal | none |
| Whole Grains | 10 | ≥1.5 oz-eq/1,000 kcal | none |
| Dairy | 10 | ≥1.3 cup-eq/1,000 kcal | none |
| Total Protein Foods | 5 | ≥2.5 oz-eq/1,000 kcal | none |
| Seafood & Plant Proteins | 5 | ≥0.8 oz-eq/1,000 kcal | none |
| Fatty Acids Ratio | 10 | (PUFA+MUFA)/SFA ≥2.5 | ≤1.2 |
| Refined Grains | 10 | ≤1.8 oz-eq/1,000 kcal | ≥4.3 oz-eq/1,000 kcal |
| Sodium | 10 | ≤1.1 g/1,000 kcal | ≥2.0 g/1,000 kcal |
| Added Sugars | 10 | ≤6.5% energy | ≥26% energy |
| Saturated Fats | 10 | ≤8% energy | ≥16% energy |
Scoring between the two standards is proportional (linear).
Predictive strength (cohort). NIH-AARP Diet and Health Study, n = 492,823, 15 y follow-up, 86,419 deaths (Reedy et al., J Nutr 2014;144(6):881-9, PMID 24572039, fetched). All-cause mortality, highest vs lowest quintile:
- HEI-2010 — men HR 0.78 (0.76–0.80); women HR 0.77 (0.74–0.80)
- AHEI-2010 — men HR 0.76 (0.74–0.78); women HR 0.76 (0.74–0.79)
- aMED — men HR 0.77 (0.75–0.79); women HR 0.76 (0.73–0.79)
- DASH — men HR 0.83 (0.80–0.85); women HR 0.78 (0.75–0.81)
This is the most decision-relevant single result in the whole literature review for us. Four structurally different indices — a US-guideline index, a Harvard index, a Mediterranean index and a BP-lowering index — produce essentially identical hazard ratios (~0.76–0.83) in the same cohort. That tells us:
- The specific index matters far less than whether it captures the same handful of underlying signals (more plants/fibre/unsaturated fat, less sodium/added sugar/refined starch/processed meat).
- There is no “best” validated index to adopt. Chasing one is chasing noise.
- A custom score that captures those same signals should be expected to perform comparably. We are not obviously reinventing something better; we are re-instantiating a well-trodden family.
Per-meal usable? Arithmetically yes, semantically no. Not recommended as our primary. See §2.1.
2.3 AHEI-2010
Harvard index, 11 components scored 0–10 each (total 110), designed explicitly to predict chronic disease rather than to measure guideline conformance. Components include vegetables, fruit, whole grains, SSBs/fruit juice, nuts/legumes, red/processed meat, trans fat, long-chain omega-3s, PUFA, sodium, and alcohol.
Predictive strength: essentially identical to HEI in the NIH-AARP head-to-head above (HR 0.76 both sexes). The Chiuve et al. 2012 J Nutr paper is titled “Alternative dietary indices both strongly predict risk of chronic disease” (PMID 22513989, citation verified; PubMed carries no abstract for it, so I did not verify its internal effect sizes in this session — [UNVERIFIED] for the specific NHS/HPFS numbers).
Per-meal usable? No. The alcohol component (scored on a J-shaped curve of daily drinks) and the trans-fat-as-%-of-long-term-energy component have no meaning for one meal. Do not use.
2.4 DASH — the trial vs the score (do not conflate)
These are two different things and the confusion matters for how hard we penalise sodium.
(a) The DASH trial (Appel et al., NEJM 1997;336:1117-24, n=459, PMID 9099655, fetched). Feeding trial, sodium intake and body weight held constant. The combination diet (fruits, vegetables, low-fat dairy, reduced saturated and total fat) vs control:
- Whole sample: −5.5 mmHg systolic / −3.0 mmHg diastolic
- Hypertensive subgroup (n=133): −11.4 / −5.5 mmHg
- Normotensive subgroup (n=326): −3.5 / −2.1 mmHg
- Fruits-and-vegetables-only diet: −2.8 systolic (−1.1 diastolic, p=0.07)
Note what this shows: most of the DASH BP effect is achievable with zero sodium reduction. Food-pattern composition, not salt, delivered 5.5 mmHg.
(b) DASH-Sodium (Sacks et al., NEJM 2001;344:3-10, n=412, PMID 11136953, fetched). Three sodium levels × 30 days each, within each diet:
- High → intermediate sodium: −2.1 mmHg systolic on control diet, −1.3 mmHg on DASH diet
- Intermediate → low sodium: a further −4.6 mmHg on control, −1.7 mmHg on DASH
- DASH + low sodium vs control + high sodium: −7.1 mmHg in normotensives, −11.5 mmHg in hypertensives
Critical read for our weighting: on an already-good dietary pattern (the DASH arm), going from intermediate to low sodium bought only 1.7 mmHg. Sodium reduction has strongly diminishing returns once the rest of the pattern is good. And these effects were measured in adults aged ~48 with elevated BP — not in a adult.
(c) DASH scores (Fung / Mellen / Dixon variants) are diet-level indices built from food-group intakes plus sodium. NIH-AARP HR 0.78–0.83 (above). Per-meal usable? No — components are daily servings of food groups.
2.5 Mediterranean diet: MDS and PREDIMED
MDS (Trichopoulou). 0–9 score. For each of 9 components, you get 1 point if your intake is above (for beneficial components) or below (for detrimental) the sex-specific median of the study population. This is a population-relative construction. A single meal has no population median. Per-meal usable? No — structurally impossible, not merely unvalidated. Any “Mediterranean score” applied to a meal is a different instrument wearing the same name.
PREDIMED (Estruch et al., NEJM 2018;378(25):e34, PMID 29897866, fetched). Get this history right:
- n = 7,447, aged 55–80, high CV risk, no CVD at baseline, Spain. Three arms: MedDiet + extra-virgin olive oil, MedDiet + mixed nuts, control (advice to reduce dietary fat). Median follow-up 4.8 y.
- The 2013 NEJM paper was withdrawn and republished in 2018. Reason: protocol deviations — enrolment of household members without randomisation, assignment without randomisation at 1 of 11 sites, and apparent inconsistent use of randomisation tables at another site. The republished analysis does not rely exclusively on the assumption that all participants were randomised.
- Republished result: 288 primary events total. EVOO arm 96 events (3.8%), nuts arm 83 (3.4%), control 109 (4.4%). HR 0.69 (95% CI 0.53–0.91) for EVOO, 0.72 (0.54–0.95) for nuts.
- Absolute effect: ~0.6–1.0 percentage points over 4.8 years in a high-risk elderly population. In a adult with near-zero 5-year CVD risk, the corresponding absolute benefit is negligible; the value of the pattern for the adult is cumulative over decades, not per-meal.
2.6 DII (Dietary Inflammatory Index)
Construction (Shivappa et al., Public Health Nutr 2014;17(8):1689-96, PMID 23941862, fetched): ~6,500 articles screened, 1,943 qualifying articles scored for whether each dietary parameter raises (+1), lowers (−1) or does not affect (0) six inflammatory biomarkers (IL-1β, IL-4, IL-6, IL-10, TNF-α, CRP). 45 food parameters. Individual intakes are expressed relative to a composite global database built from 11 food-consumption datasets from around the world — i.e. each parameter is z-scored against a world-standard mean and SD, then converted to a centred percentile.
Range on that global database: maximally pro-inflammatory +7.98, maximally anti-inflammatory −8.87, median +0.23.
Per-meal usable? No — structurally impossible. The z-score is defined against global mean daily intake. Feeding it a single meal’s nutrient totals compares one meal against a full day’s world-average intake, which drives every meal artificially “anti-inflammatory-poor”. The index has no meal-level referent, and there is none to construct without re-deriving world-standard per-meal means. Do not use.
I did not, in this session, verify the magnitude of DII’s correlation with measured CRP/IL-6 — [UNVERIFIED], though the construction being literature-derived rather than biomarker-fitted is itself a reason to expect it to be modest.
2.7 NOVA as a “scoring system”
NOVA is a 4-category nominal classification (1 unprocessed/minimally processed, 2 culinary ingredients, 3 processed, 4 ultra-processed). It is not a score:
- No intensity gradient. A meal-kit chicken breast with a stabiliser and a can of cola are both NOVA 4.
- Not ordinal in the way a score needs. NOVA 2 (oil, salt, sugar) is not “healthier than” NOVA 3.
- It is defined by formulation and ingredient provenance, not by any measurable physical property of the food.
For our purpose the operative point is: NOVA cannot rank two ready meals against each other, because essentially every meal-delivery product is NOVA 4. An input that takes the same value for ~all 800 items carries zero discriminative information. If our “processing markers” sub-score is effectively a NOVA-4 proxy, it is doing nothing but adding a constant. Check this empirically — see §5.4.
2.8 FSA/Ofcom Nutrient Profiling Model (UK)
Per-product, fully specified, and the direct ancestor of Nutri-Score. Developed 2004–05 by the BHF Health Promotion Research Group at Oxford for the FSA/Ofcom, to define foods that may not be advertised to children. All criteria per 100 g.
Validation: this is expert-judgement validation, not outcome validation. Prototypes were compared against ratings of “healthiness” for 40 foods by a survey of UK professional nutritionists; the chosen model correlated r = 0.80 (95% CI 0.73–0.86) with professional ratings. (Source: https://www.ndph.ox.ac.uk/cpnp/files/about/uk-ofcom-nutrient-profile-model.pdf, fetched.)
Algorithm (fetched verbatim):
Step 1 — A points (max 10 each; per 100 g):
| Points | Energy (kJ) | Sat fat (g) | Total sugar (g) | Sodium (mg) |
|---|---|---|---|---|
| 0 | ≤335 | ≤1 | ≤4.5 | ≤90 |
| 1 | >335 | >1 | >4.5 | >90 |
| 2 | >670 | >2 | >9 | >180 |
| 3 | >1005 | >3 | >13.5 | >270 |
| 4 | >1340 | >4 | >18 | >360 |
| 5 | >1675 | >5 | >22.5 | >450 |
| 6 | >2010 | >6 | >27 | >540 |
| 7 | >2345 | >7 | >31 | >630 |
| 8 | >2680 | >8 | >36 | >720 |
| 9 | >3015 | >9 | >40 | >810 |
| 10 | >3350 | >10 | >45 | >900 |
Step 2 — C points (max 5 each; per 100 g):
| Points | Fruit/Veg/Nuts (%) | NSP fibre (g) | or AOAC fibre (g) | Protein (g) |
|---|---|---|---|---|
| 0 | ≤40 | ≤0.7 | ≤0.9 | ≤1.6 |
| 1 | >40 | >0.7 | >0.9 | >1.6 |
| 2 | >60 | >1.4 | >1.9 | >3.2 |
| 3 | — | >2.1 | >2.8 | >4.8 |
| 4 | — | >2.8 | >3.7 | >6.4 |
| 5 | >80 | >3.5 | >4.7 | >8.0 |
Step 3 — Overall score:
- If A points < 11: score = A − C
- If A points ≥ 11 and fruit/veg/nut points = 5: score = A − C
- If A points ≥ 11 and fruit/veg/nut points < 5: score = A − (fibre points + fruit/veg/nut points) — protein points are discarded.
Cut-offs: a food is “less healthy” at score ≥ 4; a drink is “less healthy” at score ≥ 1.
Note the protein rule and why it exists: without it, salty high-protein products (processed meat, some cheeses) would earn their way out of a bad score purely on protein. This is the same design problem our score has — see §5.3.
Outcome validation of an FSA-NPM-based dietary index exists via the Nutri-Score literature (§3.3).
2.9 NRF9.3 — the closest validated analogue to what we built
This is the system our custom score most resembles, and it is worth knowing it exists.
Formula (Drewnowski; J Nutr 2009 — retrieved via search summary, formula consistent across sources, but I did not fetch the primary PDF in this session, so treat the exact denominators as [PARTIALLY VERIFIED]):
NRF9.3 = ( protein/50g + fiber/25g + vitA/5000IU + vitC/60mg + vitE/30IU
+ Ca/1000mg + Fe/18mg + Mg/400mg + K/3500mg
- satfat/20g - added_sugar/50g - sodium/2400mg ) × 100
Each nutrient is expressed as a fraction of its daily reference value, typically capped at 100% per nutrient to stop fortification gaming, and computed per 100 kcal or per RACC.
Validation: against HEI. NRF9.3 per 100 kcal explained the highest share of HEI variance among the NRF family tested — R² ≈ 0.45. Indices using both “nutrients to encourage” and “nutrients to limit” outperformed limit-only indices; performance declined when more than 6–9 encouraged nutrients were added.
Three lessons for us, and they are directly actionable:
- A capped, reference-value-normalised, per-100-kcal nutrient-density score is a legitimate, published, validated design. We are not doing something eccentric. This is the strongest single answer to “are we reinventing something?” — yes, we are reinventing NRF, and NRF is a reasonable thing to have reinvented.
- Cap each component. If our sub-scores are uncapped, a single extreme value (a very-high-protein meal) can dominate. NRF’s capping exists precisely for this.
- More components is not better. NRF performance declined past ~9 encouraged nutrients. This is direct evidence against adding more inputs to our score, and mild evidence that our 8 inputs are already at or past the useful ceiling.
2.10 FDA 2024 “healthy” claim — the only per-MEAL regulatory standard
The FDA’s updated “healthy” nutrient content claim final rule (December 2024) is food-group-based plus limits, and — uniquely among everything reviewed here — has an explicit MEAL tier. (Source: fda.gov, fetched.)
| Tier | Food-group requirement | Sat fat | Sodium | Added sugar |
|---|---|---|---|---|
| Individual food (per RACC) | e.g. ¾ oz whole grain, or ½ cup veg, or 1 oz seafood, etc. | 5–10% DV (1–2 g) | 10% DV (230 mg) | 2–10% DV (1–5 g) |
| Mixed product | ≥⅔ food-group-eq from ≥2 groups | 2 g | 230 mg | 2.5 g |
| Meal | 3 total food-group-eq from ≥3 groups | ≤4 g | ≤690 mg | ≤10 g |
This is the single most useful external benchmark in this document for our project, because it is (a) regulatory, (b) explicitly meal-level, and (c) numerically concrete. Compliance date was not captured in the fetched page — [UNVERIFIED].
Actionable: add a boolean fda_healthy_meal flag to all ~800 meals. The 690 mg sodium / 4 g sat fat / 10 g added sugar triple is a defensible, externally-sourced pass-fail line, and the fraction of the catalogue that passes is a headline number worth having.
2.11 Meal Balance Index — the one purpose-built meal-level index
Mainardi et al., Nestlé Research, PLoS ONE, December 2020 (fetched via PMC7755194). Nine nutrients: protein, total fat, fibre, potassium, calcium, iron, sodium, added sugars, saturated fat. Score 0–100, nutrient density adjusted to 2,000 kcal, each nutrient scored against a healthy range then weighted and averaged.
Applied to 147,849 meals from NHANES 2005–2014 (39,630 breakfasts, 35,687 lunches, 38,976 snacks, 33,556 dinners).
- Meals from four expert-designed exemplary dietary patterns scored 76 ± 14
- Actual NHANES participant meals scored 45 ± 14
- Correlation with HEI: Pearson r ≈ 0.6
- Higher-MBI meals contained greater densities of under-consumed micronutrients and favourable food groups not directly in the algorithm (a genuine construct-validity result)
- No association with health outcomes was reported.
Authors’ stated rationale is precisely ours: few meal-quality indicators exist, existing ones lack internal validation methodology, and “dietary advice might be more practical and easier to follow if given for meals and snacks.”
Read for us: the field’s own purpose-built meal index is (a) structurally very close to what we built — nutrient-density, ~9 nutrients, normalised to a daily energy basis, scored against ranges, weighted and averaged; (b) validated only against HEI and expert menus, not against hard outcomes; and (c) contains no processing or additive term at all. If the best published meal-level index omits processing entirely, a 24% combined processing+additive weight in ours is hard to defend on precedent as well as on evidence.
2.12 Comparison table
| System | Level | Inputs required | Validated against | Hard-outcome prediction | Per-meal usable? |
|---|---|---|---|---|---|
| HEI-2020 | Diet (day+) | Food-group equivalents + 4 nutrients | DGA conformance; construct validity | Strong (all-cause mortality HR 0.77–0.78 Q5v1, NIH-AARP n=492,823) | Arithmetically yes; semantically no — penalises focused meals |
| AHEI-2010 | Diet | Food groups + alcohol + trans fat + n-3 | Chronic disease prediction | Strong (HR 0.76) | No — alcohol & long-term trans-fat components |
| DASH score | Diet | Daily food-group servings + Na | BP trial pattern; cohorts | Strong (HR 0.78–0.83) | No |
| MDS (Trichopoulou) | Diet, population-relative | 9 components vs cohort medians | Cohort mortality | Strong (aMED HR 0.76–0.77) | No — needs a population median |
| PREDIMED/MEDAS | Diet (behavioural) | 14 adherence questions | RCT: HR 0.69/0.72, abs. 3.4–3.8% vs 4.4% over 4.8 y | Strongest causal evidence for a pattern | No — it is an adherence questionnaire |
| DII | Diet, global-relative | 45 parameters + world means/SDs | Inflammatory biomarkers (literature-derived) | Weak/modest | No — z-scored to world daily intake |
| NOVA | Product, nominal | Ingredient list + provenance | n/a (a classification, not a score) | Association only | No — no gradient; ~all our items are NOVA 4 |
| Nutri-Score 2023 | Product (100 g) | Energy, sugar, satfat, salt, protein, fibre, FVL% | Expert judgement + prospective cohorts | Modest but real (EPIC cancer HR 1.07) | Yes (with 100 g-basis caveats) |
| FSA/Ofcom NPM | Product (100 g) | Same 7 inputs | Nutritionist ratings r=0.80 (0.73–0.86) | Via Nutri-Score literature | Yes |
| NRF9.3 | Food (per 100 kcal / RACC) | 9 nutrients + 3 limits, DV-normalised | HEI (R²≈0.45) | Indirect only | Yes |
| FDA “healthy” 2024 | Product AND MEAL | Food-group eq + satfat/Na/added sugar | Regulatory/DGA | n/a (pass-fail) | Yes — has an explicit meal tier |
| Meal Balance Index | Meal | 9 nutrients @2,000 kcal density | HEI (r≈0.6) + expert menus, n=147,849 meals | None reported | Yes — purpose-built |
2.13 Verdict on §2
No single validated system is best for ranking prepared meals, and adopting one would be a downgrade.
- The strongly outcome-validated systems (HEI/AHEI/DASH/MDS) are all diet-level and mostly meal-incoherent.
- The per-item systems (Nutri-Score, FSA-NPM, NRF9.3) are usable but have weaker outcome validation and are 100 g- or 100 kcal-normalised, which is not the same as ranking meals as eaten.
- The one purpose-built meal index (MBI) has no hard-outcome validation at all and looks a lot like ours.
- The NIH-AARP head-to-head (§2.2) shows that structurally different indices capturing the same underlying signals converge on the same hazard ratios. Index choice is second-order. Signal choice is first-order.
Therefore: keep the custom score. Add Nutri-Score 2023, FSA-NPM score, and the FDA healthy-meal boolean as computed reference columns. They are cheap (same seven nutrients we already have), they give external defensibility, and disagreements flag meals worth manual review.
3. Nutri-Score 2023 — implementable specification
All tables below were extracted verbatim from the official reports of the European Scientific Committee in charge of updating the Nutri-Score:
- Solid foods + fats/oils/nuts/seeds + meat rules: “Update of the Nutri-Score algorithm” (main algorithm report, 2022) — fetched from https://mpc.gouvernement.lu/dam-assets/le-ministère/consodur/2022-main-algorithm-report-update-final.pdf
- Beverages: “Update of the Nutri-Score algorithm for beverages” (2023, C. Julia et al.) — fetched from https://www.bmleh.de/SharedDocs/Downloads/DE/_Ernaehrung/Lebensmittel-Kennzeichnung/nutri-score-update-algorithm-beverages.pdf
The new algorithm applies to products newly placed on the EU market from 31 December 2023, with a transition period to end-2025 for products already on the market.
Warning about third-party sources. During this research I fetched a commercial Nutri-Score calculator vendor’s “technical reference” page which published incorrect fibre thresholds (it gave 3.0–4.9 → 1 pt, 5.0–6.4 → 2 pt, 6.5–7.4 → 3 pt, max 5; the official table is >3.0 → 1, >4.1 → 2, >5.2 → 3, >6.3 → 4, >7.4 → 5). Implement from the official tables below, not from calculator write-ups.
3.1 Product categorisation (do this first)
Every product falls in exactly one of four buckets, each with its own tables and cut-offs:
- General foods (default — this is where all our ready meals go)
- Fats, oils, nuts and seeds — plant/animal fats and oils incl. cream, margarine, butter, oils; plus nuts (HS 0801, 0802), processed nuts (200811, 200819), ground nuts (1202), seeds (1204 linseed, 1206 sunflower, 1207 other). Chestnuts are excluded.
- Beverages — all non-alcoholic drinks incl. water, juices, nectars, smoothies, coffee/tea, and (new in 2023) milk, milk-based and fermented-milk drinks and plant-based milk analogues. Alcohol >1.2% is out of scope entirely.
- Cheese and red meat are general foods with one special rule each (§3.2.3, §3.2.4).
3.2 General foods (use this for ready meals)
3.2.1 Unfavourable (“A”) points, per 100 g
| Points | Energy (kJ) | Sugars (g) | Saturates (g) | Salt (g) |
|---|---|---|---|---|
| 0 | < 335 | ≤ 3.4 | ≤ 1 | ≤ 0.2 |
| 1 | > 335 | > 3.4 | > 1 | > 0.2 |
| 2 | > 670 | > 6.8 | > 2 | > 0.4 |
| 3 | > 1005 | > 10 | > 3 | > 0.6 |
| 4 | > 1340 | > 14 | > 4 | > 0.8 |
| 5 | > 1675 | > 17 | > 5 | > 1.0 |
| 6 | > 2010 | > 20 | > 6 | > 1.2 |
| 7 | > 2345 | > 24 | > 7 | > 1.4 |
| 8 | > 2680 | > 27 | > 8 | > 1.6 |
| 9 | > 3015 | > 31 | > 9 | > 1.8 |
| 10 | > 3350 | > 34 | > 10 | > 2.0 |
| 11 | > 37 | > 2.2 | ||
| 12 | > 41 | > 2.4 | ||
| 13 | > 44 | > 2.6 | ||
| 14 | > 48 | > 2.8 | ||
| 15 | > 51 | > 3.0 | ||
| 16 | > 3.2 | |||
| 17 | > 3.4 | |||
| 18 | > 3.6 | |||
| 19 | > 3.8 | |||
| 20 | > 4.0 |
Max A = 10 (energy) + 15 (sugars) + 10 (saturates) + 20 (salt) = 55.
Note salt in grams of NaCl, not sodium. salt_g = sodium_mg × 2.5 / 1000.
3.2.2 Favourable (“C”) points, per 100 g
| Points | Protein (g) | Fibre (g) | Fruit, vegetables & legumes (%) |
|---|---|---|---|
| 0 | ≤ 2.4 | ≤ 3.0 | ≤ 40 |
| 1 | > 2.4 | > 3.0 | > 40 |
| 2 | > 4.8 | > 4.1 | > 60 |
| 3 | > 7.2 | > 5.2 | — |
| 4 | > 9.6 | > 6.3 | — |
| 5 | > 12 | > 7.4 | > 80 |
| 6 | > 14 | ||
| 7 | > 17 |
Max C = 7 + 5 + 5 = 17.
FVL component composition (2023 change). Qualifying ingredients are Eurocodes 8.10, 8.15, 8.20, 8.25, 8.30, 8.38, 8.40, 8.42, 8.45, 8.50, 8.55, 8.60 (vegetables); 9.10, 9.20, 9.25, 9.30, 9.40, 9.50, 9.60 (fruits); 7.10 (pulses). Nuts and oils no longer count toward FVL in the general-foods algorithm (they did before 2023).
3.2.3 Computation
A = pts(energy) + pts(sugars) + pts(saturates) + pts(salt)
C = pts(protein) + pts(fibre) + pts(FVL)
if A >= 11 and product is NOT cheese:
score = A - (pts(fibre) + pts(FVL)) # protein points DISCARDED
else:
score = A - C
Cheese exception: for cheese, protein points are always counted (score = A − C) regardless of A, to avoid over-penalising naturally saturated-fat- and salt-rich dairy.
3.2.4 Red meat rule
For red meat products, protein points are capped at 2 (instead of 7). Qualifying: beef, veal, swine (pork), lamb, and also game/venison, horse, donkey, goat, camel, kangaroo.
3.2.5 Letter grades — general foods, cheese, red meat
| Score | Class | Colour |
|---|---|---|
| min to 0 | A | dark green |
| 1 to 2 | B | light green |
| 3 to 10 | C | yellow |
| 11 to 18 | D | light orange |
| 19 to max | E | dark orange |
3.3 Fats, oils, nuts and seeds
A points, per 100 g
| Points | Energy from saturates (kJ)* | Sugars (g) | Saturates/total lipids (%) | Salt (g) |
|---|---|---|---|---|
| 0 | ≤ 120 | ≤ 3.4 | < 10 | ≤ 0.2 |
| 1 | > 120 | > 3.4 | < 16 | > 0.2 |
| 2 | > 240 | > 6.8 | < 22 | > 0.4 |
| 3 | > 360 | > 10 | < 28 | > 0.6 |
| 4 | > 480 | > 14 | < 34 | > 0.8 |
| 5 | > 600 | > 17 | < 40 | > 1.0 |
| 6 | > 720 | > 20 | < 46 | > 1.2 |
| 7 | > 840 | > 24 | < 52 | > 1.4 |
| 8 | > 960 | > 27 | < 58 | > 1.6 |
| 9 | > 1080 | > 31 | < 64 | > 1.8 |
| 10 | > 1200 | > 34 | ≥ 64 | > 2.0 |
| 11–15 | > 37 / > 41 / > 44 / > 48 / > 51 | > 2.2 … > 3.0 | ||
| 16–20 | > 3.2 … > 4.0 |
* Energy from saturates (kJ/100g) = saturates (g/100g) × 37
Note the two structural differences vs general foods: energy is replaced by energy from saturates, and the saturates axis becomes a ratio of saturates to total lipids (so an oil is judged on fat quality, not fat quantity). This is why the 2023 update improved olive oil and nuts.
C points
Same tables as general foods (§3.2.2). Additionally, oils derived from FVL-qualifying ingredients do count toward FVL in this category (e.g. olive, avocado).
Computation and grades
if A >= 7: score = A - (pts(fibre) + pts(FVL))
else: score = A - C
| Score | Class |
|---|---|
| min to −6 | A |
| −5 to 2 | B |
| 3 to 10 | C |
| 11 to 18 | D |
| 19 to max | E |
3.4 Beverages
A points, per 100 mL
| Points | Energy (kJ) | Sugars (g) | Saturates (g) | Salt (g) | Non-nutritive sweeteners |
|---|---|---|---|---|---|
| 0 | ≤ 30 | ≤ 0.5 | ≤ 1 | ≤ 0.2 | absent |
| 1 | ≤ 90 | ≤ 2 | > 1 | > 0.2 | |
| 2 | ≤ 150 | ≤ 3.5 | > 2 | > 0.4 | |
| 3 | ≤ 210 | ≤ 5 | > 3 | > 0.6 | |
| 4 | ≤ 240 | ≤ 6 | > 4 | > 0.8 | presence → 4 pts |
| 5 | ≤ 270 | ≤ 7 | > 5 | > 1.0 | |
| 6 | ≤ 300 | ≤ 8 | > 6 | > 1.2 | |
| 7 | ≤ 330 | ≤ 9 | > 7 | > 1.4 | |
| 8 | ≤ 360 | ≤ 10 | > 8 | > 1.6 | |
| 9 | ≤ 390 | ≤ 11 | > 9 | > 1.8 | |
| 10 | > 390 | > 11 | > 10 | > 2.0 | |
| 11–20 | > 2.2 … > 4.0 |
The non-nutritive sweetener component is new in 2023 and is the only place in the whole Nutri-Score system where an additive is penalised — a flat +4 A-points for presence, no dose response.
C points, per 100 mL
| Points | Protein (g) | Fibre (g) | FVL (%) |
|---|---|---|---|
| 0 | ≤ 1.2 | ≤ 3 | ≤ 40 |
| 1 | > 1.2 | > 3 | — |
| 2 | > 1.5 | > 4.1 | > 40 |
| 3 | > 1.8 | > 5.2 | — |
| 4 | > 2.1 | > 6.3 | > 60 |
| 5 | > 2.4 | > 7.4 | — |
| 6 | > 2.7 | > 80 | |
| 7 | > 3.0 |
Computation and grades
score = A − C always (no A≥11 protein-discard rule for beverages).
| Score | Class |
|---|---|
| water | A (dedicated flag) |
| min to 2 | B |
| 3 to 6 | C |
| 7 to 9 | D |
| 10 to max | E |
Only water can be graded A. Every other beverage is B–E by construction.
3.5 Known failure modes — read before using it on ready meals
- 100 g basis, portion-blind. Nutri-Score judges concentration, not dose. A 200 g meal and a 500 g meal with identical per-100-g composition get identical letters despite a 2.5× difference in everything consumed. For our use case this is the single biggest problem, because meal-delivery portions vary enormously and the adult’s actual decision is “which meal should I eat”, not “which formulation is denser in badness”. Our per-meal score should NOT copy this.
- Sat-fat and salt dominate savoury protein foods. Cheese and processed meats land D/E chiefly on saturated fat, salt and energy density. The cheese exception and red-meat protein cap are explicit acknowledgements that the base algorithm mis-ranks these.
- Olive oil / nuts. Historically criticised; the 2023 fats-and-oils algorithm (saturates-to-lipid ratio, energy-from-saturates) improved these substantially. Some criticism persists that with a fixed 100 mL basis, olive oil can still only reach C even though it is used in small amounts.
- Total sugars, not added sugars. Consequence: plain yogurt and whole fruit purée are penalised for intrinsic sugars, while the algorithm cannot distinguish them from added sucrose. Our score using added sugar is a genuine improvement over Nutri-Score here, and we should say so.
- No processing, no additives (except beverage NNS). Explicitly out of scope. Nutri-Score’s own creator (Hercberg) has since ~2021 proposed a complementary, separate label for processing/additives rather than folding it into the score. That architectural choice — keep processing as a separate signal, not a weighted term inside a nutrient score — is directly applicable to us. See §5.4.
- No micronutrients. Protein is used as a proxy for micronutrient density (an explicit design decision inherited from FSA-NPM). NRF9.3 does this better if micronutrient data is available.
- No guidance for composite dishes / ready meals as a distinct class. Ready meals are simply “general foods”. There is no meal-specific Nutri-Score variant.
[Not found in this session — if one exists I did not locate it.]
3.6 Outcome validation of Nutri-Score / FSAm-NPS
The strongest prospective evidence: Deschasaux et al., PLOS Medicine 2018 — “Nutritional quality of food as represented by the FSAm-NPS nutrient profiling system underlying the Nutri-Score label and cancer risk in Europe: results from the EPIC prospective cohort study” (fetched from journals.plos.org, DOI 10.1371/journal.pmed.1002651).
- Cohort, n = 471,495 adults, 10 European countries, median follow-up 15.3 y, 49,794 incident cancers.
- Total cancer, Q5 (worst nutritional quality) vs Q1: HR 1.07 (95% CI 1.03–1.10), p-trend < 0.001.
- Absolute rates: 81.4 vs 69.5 cases per 10,000 person-years.
- Site-specific: colorectal HR 1.11 (1.01–1.22); lung in men 1.26 (1.06–1.51); post-menopausal breast 1.08 (1.00–1.16); liver in women 2.33 (1.23–4.43).
How to read this honestly. This is the best outcome validation any per-product nutrient profiling model has, and it is a 7% relative increase, ~12 extra cancers per 10,000 person-years, comparing the worst fifth to the best fifth of European diets. It is a real signal and it justifies using the model. It is also small, observational, and vulnerable to healthy-user confounding. It does not license strong claims that a one-letter Nutri-Score difference between two ready meals will change anyone’s health.
4. Weighting — is our ~76 / 10 / 14 split defensible?
4.1 The verdict, stated plainly
No. The nutrient share is roughly right; the additive share is the weakest-evidenced 14% in the model and should be cut to low single digits; the processing share is probably measuring nothing.
Specifically:
- Additives at 14% is not defensible on human outcome evidence. Reviewed below, item by item. There is no additive class with RCT hard-outcome evidence at realistic dietary exposure. Assigning additives more weight than saturated fat — which has a Cochrane-graded moderate-certainty RCT effect on CVD events — inverts the evidence hierarchy.
- Processing at 10% is probably a near-constant, not a signal. Every meal-delivery product is NOVA 4. If our processing markers are ingredient count / refined-grain presence / additive presence, they are (a) low-variance across our catalogue and (b) already correlated with nutrients we score directly. Either it adds nothing or it double-counts.
- Within the 76%, the internal balance likely over-weights saturated fat and under-weights fibre and energy density.
4.2 Effect sizes on the same footing (all verified this session unless tagged)
| Component | Best evidence | Design | Effect size | Absolute magnitude |
|---|---|---|---|---|
| Fibre | Reynolds et al., Lancet 2019;393:434-45 (WHO-commissioned). 185 prospective studies, ~135M person-years; 58 RCTs, 4,635 participants | Cohort + RCT | 15–30% lower all-cause & CV mortality, CHD, stroke, T2D, colorectal cancer, highest vs lowest fibre. Optimum 25–29 g/day, dose-response continuing above. GRADE moderate for fibre | RCTs: significantly lower body weight, systolic BP and total cholesterol |
| Saturated fat | Hooper et al., Cochrane 2020, CD011737.pub2. 15 RCTs, 16 comparisons, ~59,000 participants, ≥24 months | RCT meta-analysis | Combined CV events RR 0.79 (0.66–0.93), GRADE moderate. All-cause mortality RR 0.96 (0.90–1.03) — null. CV mortality RR 0.95 (0.80–1.12) — null. Non-fatal MI 0.97 (0.87–1.07) — null | NNTB 56 over ~4 y (primary prevention); 32 (secondary). Replacement with PUFA or carbohydrate both worked |
| Sodium | Sacks et al., NEJM 2001;344:3-10 (DASH-Sodium, n=412) | RCT | On the DASH diet: high→intermediate −1.3 mmHg; intermediate→low a further −1.7 mmHg. On the control diet: −2.1 then −4.6 mmHg | Total ~3 mmHg from a large sodium cut if the rest of the diet is already good; ~6.7 mmHg if it is not |
| Sodium (hard outcomes) | Neal et al., NEJM 2021;385:1067-77 (SSaSS, n=20,995, cluster-randomised, 4.74 y) | RCT | Stroke RR 0.86 (0.77–0.96); MACE 0.87 (0.80–0.94); death 0.88 (0.82–0.95) | Stroke 29.14 vs 33.65 per 1,000 person-years (~4.5 fewer/1,000/y). Population: mean age 65.4, 72.6% prior stroke, 88.4% hypertensive |
| Added sugar | Te Morenga et al., BMJ 2012;346:e7492. 30 RCTs + 38 cohorts | RCT + cohort | Ad libitum reduction: −0.80 kg (0.39–1.21); increase: +0.75 kg (0.30–1.19). Isocaloric exchange with other carbohydrate: +0.04 kg (−0.04 to 0.13) — NULL | SSB highest vs lowest intake, overweight/obesity OR 1.55 (1.32–1.82) |
| Protein | Morton et al., BJSM 2018;52:376-84. 49 RCTs, 1,863 participants | RCT meta-analysis | FFM +0.30 kg (0.09–0.52); 1RM +2.49 kg (0.64–4.33). No further FFM gain above ~1.62 g/kg/day | Effect larger in trained individuals (+0.75 kg, p=0.03), smaller with age |
| Refined carb / glycaemic index | Sacks et al., JAMA 2014;312:2531-41 (OmniCarb, n=163, 4 controlled-feeding diets × 5 wk) | RCT | At high carbohydrate, low-GI worsened insulin sensitivity (8.9→7.1 units, −20%, p=.002) and raised LDL (139→147 mg/dL, +6%). No benefit on BP or HDL | Reynolds 2019: GI/GL evidence graded low to very low |
| Energy density | Rolls 2006 (n=24); NIH 4-arm factorial NCT05290064 (n=36); Robinson et al. meta-analysis of 10 UPF RCTs | RCT | Rolls: −575 kcal/day for a 25% reduction in energy density, with no change in hunger ratings. NIH 4-arm: +661.6 kcal/day (525–798), p<0.0001 isolating ED alone. Robinson: intake SMD 0.71 (0.45–0.98) when UPF arms were >10% higher ED vs 0.02 (−0.51, 0.55) when ED was matched | ~70% of the entire 948 kcal/day “UPF effect” in the factorial trial. The single largest per-meal lever on energy intake |
| Overall nutrient profile (FSAm-NPS) | Deschasaux et al., PLOS Med 2018, EPIC n=471,495, 15.3 y, 49,794 cancers | Cohort | Total cancer HR 1.07 (1.03–1.10) Q5 vs Q1 | 81.4 vs 69.5 cases/10,000 person-years (~12 extra per 10,000 py) |
| Whole dietary pattern | PREDIMED (n=7,447, 4.8 y) | RCT | MACE HR 0.69 (0.53–0.91) EVOO, 0.72 (0.54–0.95) nuts | 3.8% / 3.4% vs 4.4% — ~0.6–1.0 percentage points over 4.8 y in high-risk 55–80-year-olds |
4.3 What this table actually tells us
(a) Energy density is the single biggest lever on the outcome the adult actually cares about, and it is now directly quantified. The NIH 4-arm factorial (§1.3) decomposes the UPF overeating effect into energy density (+662 kcal/day), hyperpalatability (+158) and processing-per-se (+128, NS). Robinson’s meta-analysis independently shows the UPF intake effect disappears entirely when energy density is matched (SMD 0.02, CI crossing zero). Nothing else in this table moves daily energy intake by 600+ kcal. Energy density should be a top-two weight.
(b) Fibre is the most under-rated input and should carry more weight than it usually does. It has the widest outcome base (mortality, CHD, stroke, T2D, colorectal cancer), a clean dose-response, GRADE-moderate certainty from a WHO-commissioned review, and supporting RCT evidence on weight, BP and cholesterol. No other single component in our score has that combination. It also does double duty: fibre density is a strong proxy for the intact-food-matrix property that the “processing” penalty is groping at — which is an argument for moving weight from processing into fibre rather than keeping both.
(c) Saturated fat deserves less weight than convention gives it. Cochrane finds a real effect on combined CV events (RR 0.79) but null on all-cause mortality, CV mortality and non-fatal MI, with NNTB 56 over four years in populations far higher-risk than our user. Also: replacement with carbohydrate worked as well as PUFA, meaning the penalty is really about what displaces it, which a per-meal score cannot see.
(d) Added sugar should be weighted for energy displacement, not intrinsic harm. This is the sharpest result in the table: isocaloric exchange of sugar for other carbohydrate produced zero weight change (+0.04 kg, CI crossing zero). Sugar’s causal effect on adiposity runs through calories. A per-meal score should therefore penalise added sugar mainly insofar as it displaces protein and fibre and inflates energy density — which means added sugar’s weight should be moderate, and should not be stacked on top of an energy-density penalty that is already capturing the same thing.
(e) The refined-carb / GI axis should NOT carry independent weight. OmniCarb is a well-controlled feeding RCT in which the low-GI diet made insulin sensitivity and LDL worse at high carbohydrate content. Reynolds graded GI/GL evidence low to very low while grading fibre moderate. If our “processing markers” sub-score is substantially a refined-grain/GI penalty, it is weighting an axis that the best RCT evidence does not support. Score the fibre and the whole-grain content; do not separately score the “refined-ness”.
(f) Sodium: real, but with diminishing returns and a mis-transferred population. SSaSS is a genuine RCT with hard outcomes — but in 65-year-olds, 73% with prior stroke, 88% hypertensive. DASH-Sodium shows that on an already-good dietary pattern the intermediate→low sodium step bought 1.7 mmHg. Sodium earns a solid middle weight, not a dominant one.
4.4 Additives — going through them one at a time
This is where the 14% has to justify itself.
| Additive/exposure | Best human evidence | Verified? | Verdict for scoring |
|---|---|---|---|
| Processed meat / nitrites | IARC Monograph 114 classified processed meat Group 1 (carcinogenic to humans) for colorectal cancer — Bouvard et al., Lancet Oncol 2015;16(16):1599-600 (citation and venue verified this session; the commonly-quoted “+18% per 50 g/day” figure I did not verify here — [UNVERIFIED]) |
Partly | This is the strongest item — but it is a FOOD, not an additive. It is already captured by sodium + saturated fat + a processed-meat flag. Do not double-count it as “additive burden” |
| Emulsifiers (CMC, polysorbate-80) | NutriNet-Santé cohort analyses (CVD, cancer) and a small human CMC RCT | Not verified — two PMIDs I attempted resolved to unrelated papers, so I am not quoting numbers [UNVERIFIED] |
Cohort + small-n mechanistic. Insufficient for meaningful weight. Flag, don’t weight |
| Non-sugar sweeteners | WHO 2023 guideline recommends against use of NSS for weight control | [UNVERIFIED in this session] — but the guideline is widely reported as a conditional recommendation based on low-certainty evidence, which is the operative fact |
Low certainty by the issuing body’s own grading. Note that Nutri-Score 2023 penalises NNS presence with a flat +4 A-points in beverages only — a defensible precedent for a small, flat, category-limited penalty |
| Erythritol | Witkowski et al., Nat Med 2023;29:710-18 (verified). Discovery n=1,157; validation n=2,149 US and n=833 EU; Q4 vs Q1 MACE HR 1.80 (1.18–2.77) and 2.21 (1.20–4.07); intervention arm n=8 | Verified | Weak basis for scoring. All cohorts were patients undergoing elective cardiac evaluation — i.e. high baseline CV risk. Critically, circulating erythritol is substantially produced endogenously from glucose via the pentose phosphate pathway, so plasma erythritol is plausibly a marker of dysglycaemia rather than of dietary intake. The interventional limb is 8 people with a biomarker endpoint |
| Titanium dioxide (E171) | EFSA 2021 opinion concluded E171 can no longer be considered safe as a food additive, on genotoxicity concerns | [UNVERIFIED in this session] |
This is a hazard determination, not a measured human outcome. Banned in the EU. Reasonable as a binary exclusion flag; not as a weighted continuous penalty |
| Phosphate additives | Observational CVD/kidney associations | [UNVERIFIED] |
Insufficient |
Summary judgement on additives. Across the entire class, there is not one additive with randomised hard-outcome evidence at realistic dietary exposure. The evidence is: one IARC Group 1 food (already scored via other channels), several observational cohort signals from a single cohort (NutriNet-Santé), one biomarker study with a serious reverse-causation problem, and two regulatory hazard determinations. That evidence base cannot carry 14% of a score whose other 76% includes Cochrane-grade RCT meta-analyses.
Recommended: additives → 2–4%, and preferably restructured as flags rather than weight (see §4.6).
4.5 Processing — how much weight should “how processed is it” carry?
Pulling together §1 (UPF causal evidence), §2.7 (NOVA as a construct) and the above:
- There is real experimental evidence that ultra-processed diets increase ad libitum energy intake (Hall 2019: +508 kcal/day; Dicken 2025: ~1% body weight over 8 weeks with both arms meeting dietary guidelines). That is genuine and it is why processing should not be weighted at exactly zero.
- But the causal work is done by properties we already score directly, and this has now been measured. The NIH 4-arm factorial (n=36) decomposes the effect: energy density +662 kcal/day (p<0.0001), hyperpalatability +158 kcal/day (p=0.023), and ultra-processing per se +128 kcal/day, 95% CI −8.2 to 265.0, p = 0.065 — not significant. Robinson’s meta-analysis of 10 UPF RCTs independently finds intake SMD 0.02 (−0.51, 0.55) once energy density is matched. To the extent our score already penalises energy density and rewards protein and fibre, a separate processing term is double-counting the same causal pathway.
- NOVA cannot discriminate within our catalogue. Essentially all 800 meal-delivery items are NOVA 4. A variable with near-zero variance contributes near-zero information regardless of its weight.
- The UPF-disease association attenuates substantially on nutrient adjustment (§1) — it is not cleanly independent of salt/sugar/saturated fat/fibre.
- Subgroup heterogeneity is decision-critical: the cohort signal is concentrated in processed meats and sugary drinks, while UPF breads, cereals and yogurts are neutral or inversely associated (§1). A high-protein ready meal is in the neutral part of the distribution. Applying a flat UPF penalty to it is applying an average that does not describe it.
- Nutri-Score’s own designers deliberately kept processing out of the score and proposed a separate complementary label instead. The purpose-built Meal Balance Index has no processing term at all.
Recommended: processing → 5–8%, and narrowly targeted at the sub-signals that actually carry the cohort risk (processed-meat content, sugar-sweetened liquid content), not at a generic NOVA-4 or ingredient-count marker.
4.6 Architectural recommendation: split the score
The single highest-value change is structural rather than numerical.
Stop folding processing and additives into the weighted sum. Emit them as a separate, parallel signal.
meal_output = {
"nutrition_score": 0-100, # nutrients only, the 8 nutrient inputs
"processing_flags": [...], # processed meat, SSB content, NNS, E171, emulsifiers
"fda_healthy_meal": bool, # ≤4g satfat, ≤690mg Na, ≤10g added sugar, 3 FG-eq from ≥3 groups
"nutri_score_2023": "A".."E",
"fsa_npm_score": int,
}
Why this is better than any re-weighting:
- It matches how the field’s own leaders have resolved the question (Hercberg’s complementary-label proposal; Nutri-Score, FSA-NPM, NRF9.3 and MBI all keep processing out of the score).
- It stops a weak-evidence input from silently moving the rank order of meals whose nutrient profiles differ meaningfully.
- It preserves the information — a user who cares about emulsifiers can filter on the flag — without pretending we can quantify its magnitude relative to sodium.
- It makes the score auditable: when someone asks “why did this meal score 62?”, the answer is entirely about nutrients, which is the part we can defend with RCT evidence.
If a single number is required, apply processing/additives as a small bounded penalty (cap the total at ~10 points out of 100) applied after the nutrient score, rather than as weighted terms inside it.
4.7 Recommended weighting table
Within the nutrient component. “General” = healthy adult population; “The adult” = §5.
| Input | General | The adult (29M active adult) | Justification |
|---|---|---|---|
| Fibre density | 20% | 18% | Widest outcome base (mortality, CHD, stroke, T2D, CRC), clean dose-response, GRADE moderate, plus RCT support on weight/BP/cholesterol. Also proxies food-matrix intactness, absorbing part of what “processing” was doing |
| Protein density | 12% | 24% | For the general population protein is not a strong hard-outcome predictor; its value is satiety and lean mass. For a lifting active adult in a deficit it is the primary determinant of whether lost weight is fat or muscle (Morton: plateau ~1.62 g/kg/d) |
| Energy density | 22% | 16% | Now directly quantified: +662 kcal/day in the NIH 4-arm factorial, ~70% of the whole UPF effect; and the UPF intake effect vanishes when ED is matched (SMD 0.02). Rolls: −575 kcal/day for a 25% ED cut with no increase in hunger. Down-weighted for the adult only because the adult must hit [calorie target removed] and must not be pushed toward under-fuelling — see §5.4 |
| Sodium density | 16% | 12% | RCT-backed (SSaSS, DASH-Sodium) but with strongly diminishing returns on an already-good pattern (1.7 mmHg for the int→low step) and a study population 35 years older and mostly hypertensive |
| Added sugar | 12% | 12% | Weight it for energy displacement, not toxicity: isocaloric exchange was null (+0.04 kg, CI crosses zero). Keep, but do not stack it on the energy-density penalty |
| Saturated fat | 10% | 10% | Real RCT effect on CV events (RR 0.79) but null on all-cause and CV mortality, NNTB 56 over 4 y in much higher-risk populations. LDL exposure is cumulative from young adulthood, so keep a moderate weight — but it does not deserve to outrank fibre |
| Micronutrient density (new — optional) | 8% | 8% | NRF9.3 precedent: K, Ca, Fe, Mg, vitamins A/C/E. Only add if the data exists; NRF evidence says performance declines past ~9 encouraged nutrients, so do not go further |
| Subtotal — nutrients | 100% | 100% |
And at the top level:
| Block | Current | Recommended | Change |
|---|---|---|---|
| Nutrients | ~76% | 88–92% | ↑ |
| Processing | ~10% | 5–8%, narrowly targeted at processed meat + liquid sugar | ↓ and re-scoped |
| Additives | ~14% | 2–4%, preferably as flags not weight | ↓↓ — this is the main correction |
Every component should be capped (NRF9.3 precedent) so no single input can dominate, and scored on per-meal-as-served quantities plus per-1,000-kcal density, not per 100 g (Nutri-Score’s portion-blindness is its worst property for our use case, and we should not inherit it).
5. Does the general-population evidence transfer to the adult?
Subject the scores are calibrated for: a healthy general-population adult. [personal profile and health goals removed]
5.1 The transfer problem, stated honestly
Look at the populations that generated the evidence in §4.2:
| Trial | Population | Age |
|---|---|---|
| DASH (n=459) | BP 80–95 diastolic, i.e. high-normal to stage-1 | adults, mean ~46 |
| DASH-Sodium (n=412) | as above | adults |
| SSaSS (n=20,995) | 72.6% prior stroke, 88.4% hypertensive | mean 65.4 |
| PREDIMED (n=7,447) | high cardiovascular risk | 55–80 |
| Cochrane sat-fat (15 RCTs, ~59,000) | mostly middle-aged, many with or at risk of CVD | middle-aged+ |
| EPIC/FSAm-NPS (n=471,495) | general European adults | mostly middle-aged |
Not one of these studied healthy adults. The adult’s absolute 5- and 10-year risk of the hard outcomes these trials measured is close to zero. The relative risks transfer reasonably (biology is broadly similar); the absolute benefits do not, because absolute benefit = relative reduction × baseline risk, and the adult’s baseline risk is tiny.
The honest framing: for the adult, the CVD/cancer-mediated components of the score are a 30-year bet, not a near-term intervention. They still matter — atherosclerosis and BP trajectories are cumulative from the twenties, and the habits formed now determine the exposure — but they should not out-weigh the components that affect the adult’s stated goals this year.
5.2 What transfers strongly
Protein. Morton et al. (49 RCTs, 1,863 participants) is the one piece of evidence in this whole review whose population actually resembles the adult: healthy adults doing resistance training. Its findings apply directly:
- Supplemental protein increased FFM by 0.30 kg (0.09–0.52) and 1RM by 2.49 kg (0.64–4.33)
- The effect was larger in resistance-trained individuals (+0.75 kg, p=0.03) — i.e. larger for the adult than for the average participant
- No further FFM benefit above ~1.62 g/kg/day
Two implications: (a) protein should carry the largest single weight in the adult’s version of the score, because in an energy deficit it determines whether the weight the adult loses is fat or muscle; (b) there is a defensible ceiling. Once a meal’s protein pushes the adult’s daily total past ~1.6–1.8 g/kg, additional protein density stops earning points. Our score should cap protein reward, not reward it monotonically — otherwise it will rank a 70 g-protein meal above a 45 g-protein meal with much better fibre, which is wrong for the adult.
Fibre and energy density. Both act through satiety and energy intake, which is the actual causal path to visceral fat. These transfer well because the mechanism is not disease-risk-mediated.
5.3 Where I think the adult is wrong: sodium
The adult wants to reduce sodium. Three reasons to weight this lower for the adult than for the general population:
- Diminishing returns on a good diet. DASH-Sodium: on the DASH pattern, cutting from intermediate to low sodium bought 1.7 mmHg systolic. The adult’s meals are already going to be scored toward a good pattern. The marginal mmHg is small.
- The hard-outcome evidence is from a population unlike the adult. SSaSS’s 4.5-fewer-strokes-per-1,000-person-years came in 65-year-olds with prior stroke and hypertension. There is no equivalent trial in normotensive young athletes, and the absolute benefit in that group is necessarily far smaller.
- The adult is an active adult and sodium is a performance and safety nutrient. Sweat sodium losses during hard training, especially in heat, are substantial and vary widely between individuals. Aggressively minimising dietary sodium while training hard risks impaired plasma volume, cramping, and in extreme cases exercise-associated hyponatraemia.
[UNVERIFIED — I did not verify quantitative sweat-sodium loss figures in this session. This is a directional argument from established sports-nutrition principle, not a measured number, and it should be checked before acting on it.]
Recommendation: keep a moderate sodium weight (long-run BP trajectory is a real, cumulative consideration and there is no downside to preferring the lower-sodium of two otherwise-equal meals), but do not let sodium be a top-two weight, and do not let a hard sodium ceiling knock out high-protein meals. If the adult tracks BP and it is normal, that is the measurement that should govern this — not a population guideline written for sedentary hypertension-prone adults.
5.4 Where the score could actively harm the adult: energy density
This is the most important practical point in §5.
The adult needs ~[calorie target removed]/day. Across 3–4 meals that is roughly 700–950 kcal per meal. A score that treats energy density as monotonically “lower is better” will systematically rank bulky, low-calorie, high-water meals at the top — exactly the meals that make it hard for the adult to hit [calorie target removed], and under-fuelling degrades training adaptation, recovery and, over time, VO2max progression.
Recommendation: score energy density as distance from a target band, not monotonically. Something like: full marks in a band appropriate to a ~750–900 kcal meal at reasonable volume, with penalties both above (calorie-dense, easy to overshoot) and below (so dilute the adult cannot fuel on it). This single change probably matters more to the adult’s actual outcomes than the entire processing-and-additives debate.
5.5 VO2max — be honest that we cannot score for it
VO2max is overwhelmingly determined by training stimulus, not by meal composition. The dietary contributions are narrow:
- Body composition — VO2max is expressed per kg, so fat loss raises it arithmetically. This is already captured by the energy/protein/fibre triad.
- Adequate carbohydrate availability to support the training that drives the adaptation. Note this cuts against aggressive carbohydrate or sugar penalties around training.
- Iron status (oxygen carriage) — worth capturing if micronutrient data exists (§4.7).
We should not claim the meal score targets VO2max. Its honest contribution is: keep the adult adequately fuelled, adequately carbohydrate-supplied, iron-sufficient, and leaner. Anything beyond that is training, not diet.
5.6 Where the UPF question lands for the adult specifically
The adult’s decision is “which of these 800 ready meals should I order.” All of them are NOVA 4. The realistic worst case from the UPF literature — the Hall-type ad libitum overconsumption effect — is a weight-gain mechanism, and the adult is (a) tracking intake and (b) in a deliberate deficit. The mechanism that makes UPF risky for a free-living population that eats to appetite is substantially neutralised by someone who weighs and targets the adult’s intake.
What is left of the UPF concern for the adult is the subgroup signal (§1): processed meat and sugary drinks. Those are worth flagging specifically. A generic “this is processed” penalty applied to a high-protein chicken-and-vegetables tray is not supported by the subgroup data and will distort the adult’s rankings against the adult’s own goals.
5.7 The adult’s weighting, summarised
| Input | Weight | Note |
|---|---|---|
| Protein density | 24% | Capped at the level that puts the adult’s day at ~1.6–1.8 g/kg. Do not reward beyond |
| Fibre density | 18% | Best-evidenced input overall; also satiety |
| Energy density | 16% | Target band, not monotonic. +662 kcal/day effect in the NIH factorial makes this the biggest intake lever — but see §5.4 for why it must not be monotonic for the adult |
| Sodium density | 12% | Down-weighted vs general population; governed by the adult’s actual BP |
| Added sugar | 12% | Displacement, not toxicity. Do not penalise peri-workout carbohydrate as if it were the same thing |
| Saturated fat | 10% | Cumulative-LDL argument keeps it in; null-on-mortality keeps it modest |
| Micronutrients (incl. iron) | 8% | Optional, if data available |
| — | ||
| Processing flags | separate | Targeted: processed meat, liquid sugar |
| Additive flags | separate | Informational, unweighted |
6. Honest limitations of this review
- Nutritional epidemiology is weak evidence and this document should not be read as if it were not. Almost all the disease-outcome findings above are observational, FFQ-based, and subject to (a) healthy-user confounding — people who eat less UPF also smoke less, exercise more and are richer; (b) substantial measurement error, since FFQs were never designed to capture processing; and (c) reverse causation. Fang 2024’s own note that removing smoking pack-years from the model produced a “much stronger positive association” is a clean demonstration of how much confounding is in play.
- Effect sizes shrink as adjustment improves. The pattern across the UPF literature is HR 1.22 → 1.04 → null as covariates are added. That pattern is more consistent with confounding than with a strong causal effect.
- The single most decision-relevant trial (NIH 4-arm, NCT05290064) is registry-posted and NOT peer-reviewed. Its decomposition is the backbone of the §4.5 recommendation. If it fails peer review or the numbers change, revisit the processing weight.
- Robinson et al.’s ED-matched meta-analysis is a medRxiv preprint, not peer-reviewed.
- Items I could not verify in this session and have flagged inline: the per-50 g/day processed-meat RR; the NutriNet-Santé emulsifier hazard ratios (two PMIDs I tried resolved to unrelated papers); the WHO 2023 non-sugar-sweetener guideline text; the EFSA 2021 E171 opinion text; Chiuve 2012’s internal effect sizes; sweat-sodium loss figures for athletes; the FDA “healthy” rule compliance date; the exact NRF9.3 denominators.
- No Mendelian randomisation of UPF exists, so the strongest quasi-experimental tool in nutrition is unavailable for this question.
- None of the hard-outcome evidence comes from a population resembling our user. See §5.1.
7. Concrete recommendations
- Keep the custom score. Do not adopt HEI, AHEI, DASH, MDS or DII — the first is meal-incoherent in practice and the rest are meal-incoherent in principle.
- Cut the additive weight from ~14% to 2–4%, and prefer to move it out of the weighted sum entirely into flags. No additive class has randomised hard-outcome evidence at realistic exposure.
- Cut the processing weight from ~10% to 5–8% and re-scope it. Stop scoring “is it ultra-processed.” Score the two things the cohorts consistently incriminate: processed/cured meat content and sugar-sweetened (and artificially-sweetened) liquid content. Removing exactly those two categories nulls the entire CVD association in NHS/NHSII/HPFS.
- Do not penalise UPF breads, cereals, packaged snacks or yogurt as a class. Multiple large cohorts show null or inverse associations. Penalising them moves the score against the evidence.
- Raise the energy-density weight to ~22% (general) / ~16% (the adult) — it is now directly measured as ~70% of the UPF overeating effect. But score it as a target band, not monotonically, so the score cannot push a [calorie target removed]/day active adult toward under-fuelling.
- Raise fibre to ~18–20%. Best combination of outcome breadth, dose-response and GRADE certainty of anything in the score, and it partly absorbs what the processing term was reaching for.
- Cap every component (NRF9.3 precedent), especially protein — reward it up to the level that puts the adult’s day at ~1.6–1.8 g/kg and no further.
- Score per meal as served, plus per-1,000-kcal density. Do NOT adopt Nutri-Score’s per-100 g basis — portion-blindness is its worst property for our use case.
- Add three cheap reference columns computed from the seven nutrients we already have:
nutri_score_2023(algorithm in §3),fsa_npm_score(§2.8), andfda_healthy_mealboolean (≤4 g sat fat, ≤690 mg sodium, ≤10 g added sugar, 3 food-group-eq from ≥3 groups). Use disagreement between our score and Nutri-Score as a manual-review trigger. - Run one diagnostic before anything else: compute the variance of the processing and additive sub-scores across all 800 meals. If they are near-constant (which is likely, since essentially every item is NOVA 4), they are contributing no ranking information at any weight, and the 24% they currently hold is being wasted.
8. Primary sources verified in this session
Randomised trials
- Hall KD et al. Cell Metab 2019;30(1):67-77 (PMID 31105044); errata 2019;30(1):226 (PMID 31269427) and 2020;32(4):690 (PMID 33027677)
- Dicken SJ et al. (UPDATE) Nat Med 2025;31(10):3297-3308 (PMC12532614)
- NIH 4-arm factorial, ClinicalTrials.gov NCT05290064 — results posted 2026-08-12, not peer-reviewed
- Forde CG et al. AJCN 2026 — UPF-Slow vs UPF-Fast, n=41
- Appel LJ et al. (DASH) NEJM 1997;336:1117-24 (PMID 9099655)
- Sacks FM et al. (DASH-Sodium) NEJM 2001;344:3-10 (PMID 11136953)
- Sacks FM et al. (OmniCarb) JAMA 2014;312:2531-41 (PMID 25514303)
- Neal B et al. (SSaSS) NEJM 2021;385:1067-77 (PMID 34459569)
- Estruch R et al. (PREDIMED, republished) NEJM 2018;378(25):e34 (PMID 29897866)
- Gosby AK et al. 2011 — protein leverage RCT, n=22
- Witkowski M et al. Nat Med 2023;29:710-18 (PMID 36849732)
Systematic reviews / meta-analyses
- Reynolds A et al. Lancet 2019;393:434-45 (PMID 30638909)
- Hooper L et al. Cochrane 2020 CD011737.pub2 (PMID 32428300) — note pub3 update exists
- Morton RW et al. BJSM 2018;52:376-84 (PMID 28698222)
- Te Morenga L et al. BMJ 2012;346:e7492 (PMID 23321486)
- Lane MM et al. BMJ 2024;384:e077310 (PMC10899807)
- Robinson E et al. — ED-matched UPF RCT meta-analysis, medRxiv preprint, not peer-reviewed
Cohorts
- Fang Z et al. BMJ 2024;385:e078476 (PMC11077436) — NHS+HPFS mortality, UPF subgroups
- Mendoza K et al. Lancet Reg Health Am 2024;37:100859 (PMC11403639) — CVD, UPF subgroups
- Chen Z et al. Diabetes Care 2023;46:1335-44 (PMC10300524) — T2D
- Dicken SJ et al. Lancet Reg Health Eur 2024 (PMC11551512) — EPIC T2D
- Cordova R et al. Lancet Reg Health Eur 2023;35:100771 (PMC10730313) — EPIC multimorbidity
- Reedy J et al. J Nutr 2014;144:881-9 (PMID 24572039) — NIH-AARP, 4 indices head-to-head
- Deschasaux M et al. PLOS Med 2018, doi:10.1371/journal.pmed.1002651 — EPIC, FSAm-NPS and cancer
- Braesco V et al. Eur J Clin Nutr 2022 — NOVA inter-rater reliability, Fleiss κ 0.32-0.34
- Khandpur N et al. 2021 (PMC8453454) — NOVA coding of FFQs
Algorithms and official documents (fetched)
- Nutri-Score Scientific Committee, main algorithm update report 2022 (mpc.gouvernement.lu)
- Nutri-Score Scientific Committee, beverages algorithm update report 2023 (bmleh.de)
- UK Ofcom/FSA nutrient profiling model (ndph.ox.ac.uk)
- HEI-2020 scoring standards and research uses (epi.grants.cancer.gov/hei)
- FDA “healthy” claim final rule (fda.gov)
- Mainardi F et al., Meal Balance Index, PLoS ONE 2020 (PMC7755194)
Debate / critique
- Astrup A & Monteiro CA, “Does the concept of ultra-processed foods help inform dietary guidelines?” Am J Clin Nutr 2022;116(6):1476-81 (YES), 1482-88 (NO), 1489-91 (consensus)
- Bouvard V et al. (IARC Monograph 114) Lancet Oncol 2015;16(16):1599-600 (PMID 26514947)