Working research document, copied from the author’s notes on 2026-09-06. Rough, long, and unedited apart from removing personal details. The finding page summarizes it.

Food & ingredient evidence rubric

A reusable, evidence-weighted way to judge whether a food is actually good or bad for you — as opposed to whether it sounds clean. Built 2026-09-05 during the meal-delivery comparison, but written to be general: use it for groceries, restaurant choices, protein bars, anything with a label.

Provenance. Evidence review conducted by a Fable subagent, 2026-09-05, then implemented in a scoring script. Citations are marked [V] where the agent verified them by web search in that session, [R] where recalled from training and not re-verified. Treat [R] citations as leads, not proof — check before quoting them anywhere that matters.


Headline findings

The short version, after six deep-research passes. Ordered by how much they should change behaviour.

  1. “Ultra-processed” is largely a proxy for calorie density, but the direct effect is not excluded. The ~950 kcal/day overeating effect (NIH 4-arm factorial NCT05290064, 38 enrolled / 36 completed, one-week diets; 2026, registry-posted, not peer-reviewed) decomposes into energy density (+662, p<0.0001), hyperpalatability (+158, p=0.023), and processing per se (+128 kcal/day, 95% CI -8 to 265, p=0.065, not statistically significant). Note what that interval permits: a residual direct effect of processing up to roughly +250 kcal/day is entirely consistent with the data. Most of the UPF effect is carried by energy density and palatability, which are themselves properties that processing produces, so calling them “the real cause” and processing “just a proxy” is a distinction without much practical difference. The matched-energy-density meta-analytic SMD of 0.02 is 5 comparisons with a 95% CI of -0.51 to 0.55 (June 2026 medRxiv preprint), an interval wide enough to contain most plausible effects, not a demonstrated null.

  2. Processed meat is the clearest bad actor, and the diabetes risk exceeds the cancer risk. The ~4% -> ~4.7% lifetime colorectal cancer figure is this author’s derivation, applying IARC’s 18% relative increase per 50 g/day to a ~4% baseline lifetime risk; IARC itself publishes the 18%, not the absolute move. The type 2 diabetes relative risk spans a wide literature range: 1.15 (Li 2024) to 1.51 (Pan 2011), with Micha 2010 at 1.19. Applied to a ~40% lifetime baseline that range is a +4 to +12 point absolute move (1-(1-0.40)^RR), so the “diabetes exceeds cancer” ordering is robust across the whole range even though the magnitude is not pinned down. The commonly quoted +6 comes from multiplying the hazard ratio by the baseline instead of applying it to a cumulative risk. Processed meat also carries ~622 mg sodium per 50 g.

    “Uncured” and “cultured celery powder” deliver the same nitrite. The regulatory story is more tangled than the usual one-line version. 9 CFR 424.22 sets 156 ppm ingoing nitrite for comminuted products, 200 ppm for pumped/immersion-cured products, and 120 ppm for bacon: three ceilings, not one. Celery powder is not an approved curing agent under 9 CFR 424.21(c) at all; FSIS applies those same maxima to natural nitrite sources through a note in FSIS Directive 7120.1, which is guidance, not regulation, and the long-promised “cured/uncured” labeling rulemaking has never published. So the equivalence does not rest on a shared regulatory ceiling. It rests on measured residual nitrite: Nunez De Gonzalez 2012 assayed 470 retail products and found no difference between conventionally cured and “natural/uncured” (https://pubs.acs.org/doi/10.1021/jf204611k; FSIS petition response https://www.fsis.usda.gov/sites/default/files/media_file/2021-04/19-03-fsis-final-response.pdf). The label is marketing.

  3. Liquid sugar has the most suggestive depot-specific finding for visceral fat, but it is directional rather than clean. In Stanhope 2009 (n=32 total, fructose arm n=17, overweight/obese adults aged 40-72, 25% of energy as fructose for 10 weeks), at matched weight gain fructose loaded visceral fat while glucose loaded subcutaneous. The male-specific split (+18.1% vs -0.6%) is a post-hoc sex subgroup analysis at p=0.049 in a fructose arm of 17 people, in a population older and heavier than the adult (https://pmc.ncbi.nlm.nih.gov/articles/PMC2673878/). Treat it as a direction, not an established depot mechanism. Compensation for liquid calories is -17% versus +118% for solid.

  4. Remove processed meat and sugary drinks from the UPF variable and the cardiovascular association goes from 1.11 to 1.00. The entire cohort signal was those two categories. Ultra-processed breads and cereals are mixed by outcome, not uniformly protective: inversely associated with type 2 diabetes in EPIC (HR 0.65 per 10% of intake) and inversely with CVD and CHD in Mendoza 2024 (https://pmc.ncbi.nlm.nih.gov/articles/PMC11403639/), but ultra-processed breakfast foods are positively associated with all-cause mortality in Fang 2024 (HR 1.04, 1.02-1.07; https://pmc.ncbi.nlm.nih.gov/articles/PMC11077436/). Which grain subgroup you look at, and which endpoint, changes the sign. Ready-to-heat mixed dishes are null for mortality (HR 1.02, 0.99-1.05) and CVD — but do not over-read that null. Meat-, poultry- and seafood-based ready-to-EAT products are the strongest mortality subgroup in the same analysis (HR 1.13, 1.10-1.16), and ready meals are weak-positive for type 2 diabetes. “Ready meal” is not a clean null, and this corpus sits in exactly that contested space.

  5. Additives are the weakest-supported pathway in the field. Weakest- supported is not the same as disproven. Across the 804 of 816 real meals that publish an ingredient list, 94% score zero on additives and 89% score zero on processing. But that 94% is corpus prevalence subject to the disclosure-depth trap (see below), not a statement about evidence strength, and the two must not be conflated. The animal work runs at large multiples of estimated human intake for polysorbate-80 and carboxymethylcellulose and enormous ones for BHA (denominators spelled out in the dose-multiple notes below); the much larger multiples that circulate could not be sourced. Human trials at dietary doses are mixed, not null, and both sides of that are small:

    • Healthy volunteers. Chassaing 2022 is one 11-day feeding trial, n=16 (7 on carboxymethylcellulose). It found reduced microbial diversity, reduced short-chain fatty acids, and mucus-layer bacterial encroachment in 2 of the 7 CMC subjects, the strongest mechanistic finding in the additive literature. Eleven days is too short to expect inflammatory-marker change, so “shifts without inflammation” describes the study’s duration as much as the additive’s effect (https://pubmed.ncbi.nlm.nih.gov/34774538/). Wellens 2026 (polysorbate-80 arm, n=60) is the larger follow-on.
    • Crohn’s disease. The small four-week feeding trial (n=24 randomized) that found “no significant difference” is underpowered to detect the effect actually at issue: a 49% vs 31% response difference needs far more than 24 subjects, so its null is uninformative rather than contradictory. ADDapt (n=154, 49.4% vs 30.7% response, p=0.019) is the one adequately powered trial of emulsifier restriction, and it is positive. It remains an ECCO 2025 conference abstract, not a full peer-reviewed paper (https://academic.oup.com/ecco-jcc/article/19/Supplement_1/i262/7967009). Calling this literature “split” gives the underpowered null equal billing with the powered positive.

    The operational conclusion is about ranking information, not about safety: an input that is zero for 94% of the corpus cannot separate meals no matter how it is weighted, which argues for surfacing additives as flags rather than as scored weight. It does not argue for ignoring them.

  6. Seed oils: no harm signal, and what evidence exists leans the other way. Thirty pooled cohorts using blood fatty-acid biomarkers rather than food questionnaires find higher linoleic acid associated with LOWER cardiovascular mortality, though these are biomarker cohorts and so still observational. The randomized mortality data are null rather than protective. Inflammatory markers do not rise: 15 RCTs in Johnson & Fritsche 2012, and 30 RCTs (n=1,377) in Su 2017 with a CRP SMD of 0.09 (95% CI -0.05 to 0.24), i.e. no change (https://pubmed.ncbi.nlm.nih.gov/28752873/).

  7. Fiber has the strongest outcome evidence of anything on this list — 185 prospective studies plus 58 RCTs. The 15-30% lower all-cause and cardiovascular mortality figures come from the prospective cohorts, GRADE moderate certainty; the 58 RCTs (n=4,635) measured weight, blood pressure and total cholesterol only, and no RCT here measured mortality. The dose-response “greater benefit at higher intake” statement is for cardiovascular disease, type 2 diabetes, colorectal cancer and breast cancer (https://pubmed.ncbi.nlm.nih.gov/30638909/). Target 35-40 g/day at a [calorie target removed] kcal intake.

  8. Sodium matters, but moderately, for a normotensive active adult. The blood pressure effect is 1 to 6 mmHg systolic depending on baseline diet and the size of the sodium reduction, not a flat 1-2. DASH-Sodium: on the control (typical American) diet, high-to-low sodium lowered systolic 6.7 mmHg overall and about 5.6 mmHg in normotensives; on the DASH diet the same sodium change was worth about 3.0 mmHg (https://pubmed.ncbi.nlm.nih.gov/11136953/). Graudal 2020 gives -1.14 mmHg in normotensive whites, the low end of the range and the origin of the low-end shorthand; Filippini 2021 finds -2.3 mmHg per 100 mmol/day with a linear dose-response and no threshold (https://pmc.ncbi.nlm.nih.gov/articles/PMC8094404/). No trial has shown a mortality benefit from sodium reduction alone, but the outcome data are not empty: TOHP long-term follow-up found 25% fewer cardiovascular events (RR 0.75, 0.57-0.99) with all-cause mortality directionally lower but not significant (RR 0.80, 0.51-1.26; https://pmc.ncbi.nlm.nih.gov/articles/PMC1857760/), and SSaSS found a 12% reduction in death (RR 0.88, 0.82-0.95) with a potassium-enriched salt substitute, which changes both sodium and potassium (https://pubmed.ncbi.nlm.nih.gov/34459569/). The J-curve is most likely a measurement artifact: He 2018 shows a linear relationship using measured 24h urine in TOHP and a J-curve using Kawasaki-formula estimates from spot urine in the same data (https://academic.oup.com/ije/article/47/6/1784/5040736). The PURE investigators dispute this reading, so treat it as the better- supported side of a live argument rather than settled.

  9. Selection beats sourcing. Within-service variation in meal quality is ~1.8x between-service variation. Picking well inside a mediocre service beats picking randomly inside a good one — but a service’s ceiling is not something you can pick your way past.

  10. NOVA tells you almost nothing about a specific food. Inter-rater reliability among 159 expert raters is Fleiss kappa 0.32-0.34, “fair”, a chance-corrected statistic, not a percent-agreement figure, so reading it as a raw fraction of agreements is a mistake. The concrete numbers are starker than the kappa in one direction and gentler in another: only 3 of 120 marketed foods were placed in the same category by every rater, yet 90 of those 120 drew a 91% NOVA-4 assignment rate, so raters converge strongly on the ultra-processed bucket and disagree mainly at the boundaries (Braesco 2022, https://pmc.ncbi.nlm.nih.gov/articles/PMC9436773/). The same dietary data still yields 7.9% or 45.9% ultra-processed depending on the system (Martinez-Perez 2021, PREDIMED-Plus, Nutrients 13:2471).

The practical distillation: avoid processed meat and liquid sugar; keep calorie density down; hit protein and fiber; watch sodium loosely. Seed oils and MSG carry no penalty. Additives are a second-order concern with mixed human evidence, and because nearly every meal scores zero on them they are better surfaced as flags than folded into a ranking weight. Treat them as low priority, not as settled-safe.

The headline

Additives are a second-order concern. For a healthy adult at realistic dietary doses, the outcome-relevant levers in a packaged or prepared meal are, in descending order of evidence strength:

  1. Total energy / energy density
  2. Protein
  3. Fiber
  4. Sodium
  5. Added sugar and refined starch
  6. Saturated fat
  7. Potassium / Na:K ratio
  8. Degree of processing (as a whole, not additive-by-additive)
  9. Individual additives — below everything above

The common failure mode is inverting this: paying a premium for “no gums, no natural flavors, no seed oils” while ignoring sodium, fiber and energy density. That trade is not supported by the evidence.


Tier 1 — good human evidence of harm at realistic doses

Empty. Nothing on the usual clean-label avoid-list qualifies for a healthy adult at food doses. This is itself the finding.

Tier 2 — real but uncertain signal; worth a modest penalty

Weights below are the ones score.py actually applies. Each is a fraction of the 4-point additive block; the fractions sum and are then capped at 1.0, so no meal can lose more than 4 points here however many additives it carries. A meal that trips nothing scores 0 and 94% of meals with a published list do.

Ingredient Evidence Weight (of 1.0)
Inorganic phosphates (sodium/potassium phosphate, pyro-/tripoly-/hexameta-phosphate, phosphoric acid, disodium inosinate) The widely-repeated “~90-100% bioavailable vs ~40-60% for organic phosphorus” claim is an in-vitro digestibility ceiling, not an absorption measurement [R, Calvo & Uribarri 2013]. Mendelian randomisation finds no causal effect of serum phosphate on CHD, heart failure, atrial fibrillation or hypertension, but the same MR literature does find genetically predicted phosphate causal for valvular heart disease, so “MR is null” overstates it (https://pmc.ncbi.nlm.nih.gov/articles/PMC9579538/). Controlled feeding at realistic doses is more null than previously written here: Gutiérrez 2015 (n=10, 2 weeks) found no significant change in serum phosphorus, calcium, FGF23 or PTH; the only movers were osteopontin (+10%) and osteocalcin (+11%), bone-turnover markers of unclear outcome relevance (https://pmc.ncbi.nlm.nih.gov/articles/PMC4702463/). Acute phosphate load impaired flow-mediated dilation, n=11 [R, Shuto 2009]; Framingham Offspring n≈3,400 linked serum phosphate to CVD in non-CKD adults [R, Dhingra 2007]. EFSA 2019 set a group ADI of 40 mg/kg/day, and the exceedance is in infants, toddlers and children at mean intake and adolescents at P95, not in adults (https://www.ncbi.nlm.nih.gov/pmc/articles/PMC7009158/), so the ADI-exceedance argument does not apply to the adult at all. Tier 2 borderline for normal kidney function — kept at the top of the block on dose/mechanism alignment, not on a demonstrated outcome. Bonus signal: phosphates track cured/processed meats, so the flag doubles as a sodium proxy. APPLIED 2026-09-06: the 0.6 weight, formerly the largest in the additive block, rested on borderline evidence that got weaker on re-checking, and has been dropped to 0.25 in score.py. The enumerated compound list and the exclusion list are unchanged — only the weight moved. 0.25
Erythritol / xylitol (incl. “birch sugar”) The only sweeteners with no dose gap: the ~30 g interventional dose is, per the authors, one can of an erythritol-sweetened drink or a pint of keto ice cream. Witkowski 2024, a non-randomized parallel-group trial (n=10 per arm, 30 g erythritol vs 30 g glucose), found enhanced agonist-stimulated platelet aggregation ex vivo [R]. A mechanistic endpoint, not a clotting outcome; authors have declared conflicts and a published rebuttal exists. The observational signal is likely reverse causation; the pharmacodynamics survive that critique. Design-quality note: this row is weighted 0.5 on a non-randomized n=10/arm study, while carrageenan is weighted 0.25 on Wagner 2024, a randomized double-blind crossover (n=20, 14 days, permeability p=0.03, https://bmcmedicine.biomedcentral.com/articles/10.1186/s12916-024-03771-8), the better-designed of the two. The weights are ordered opposite to study quality. 0.5
Potassium bromate / bromated flour Renal and thyroid tumours in both sexes of rat, genotoxic in vivo, IARC Group 2B. Not permitted in the EU or UK; banned in Canada, Brazil, China and India; still permitted in the US, with California banning it from 2027. Margins are large and human data are absent entirely. The 0.5 weight is therefore a precautionary flag on a free-to-avoid ingredient, not a Tier 2 evidence rating. It does not belong in the same evidential class as rows backed by human trials, and is listed here only because the block is implemented as one table. 0.5
TBHQ NTP TR-459, the 2-year bioassay, reports no evidence of carcinogenic activity in both species and that finding stands. The allergy signal is much thinner than previously written: it is one laboratory (the Rockwell group), in a mouse food-allergy model, at 0.0014% of diet (~14 ppm). That 14 ppm is a concentration comparison against permitted food levels, not a body-weight dose comparison; converting it gives roughly 2 mg/kg/day, about 3x the 0.7 mg/kg/day ADI. No published NOAEL-to-ADI ratio and no ~7-30x margin exists. The previous “NOAEL ~1.2x the ADI, margin ~7-30x” line was not sourceable and has been removed (https://pmc.ncbi.nlm.nih.gov/articles/PMC11922944/). Implementation note: TBHQ used to be listed in both the TBHQ detector and the catch-all OTHER_ADDITIVE list, so a TBHQ meal accumulated 0.4 + 0.25 = 0.65, the largest single additive penalty the scorer could apply in practice, on single-lab animal evidence. APPLIED 2026-09-06: the weight is now 0.25 and TBHQ has been removed from OTHER_ADDITIVE, so a hit contributes 0.25 exactly once. Note that no meal in the current 816-meal corpus declares TBHQ, so this is a correctness fix for future data, not a rescoring. 0.25
Carrageenan Wagner 2024 BMC Medicine: randomized double-blind crossover, n=20 healthy young men, 250 mg 2×/day for 2 weeks, increased small-intestinal permeability (p=0.03) [V] (https://bmcmedicine.biomedcentral.com/articles/10.1186/s12916-024-03771-8). The NutriNet-Santé emulsifier signal is carrageenan-specific only for diabetes: Sellem 2023 (cardiovascular disease) found no association with carrageenans, while Salame 2024 found type 2 diabetes HR 1.03 per 100 mg/day (https://pubmed.ncbi.nlm.nih.gov/38663950/). Citing “HR ~1.05 per SD for several emulsifiers” without naming the outcome overstates the carrageenan case. Older colitis work used degraded carrageenan (poligeenan), not food-grade [R]. 0.25
Artificial sweeteners (sucralose, aspartame, acesulfame) Suez 2022 Cell, n=120, 2 weeks: sucralose/saccharin altered microbiome and glucose tolerance in some individuals [R]. RCT meta-analyses show neutral-to-modest weight benefit vs sugar [R]. Mostly a junk-food marker. 0.25
BHA / BHT Rodent forestomach tumours, in an organ humans lack. Stating the denominator: the rat effect dose is 2% of diet, roughly 1,000 mg/kg bw/day. EFSA 2011 estimates adult mean intake at 0.03-0.12 mg/kg bw/day (a ratio of ~8,000-33,000x) and adult P95 at 0.08-1.12 mg/kg bw/day (~900-12,500x) (https://efsa.onlinelibrary.wiley.com/doi/abs/10.2903/j.efsa.2011.2392). A bare multiple with no denominator is uninterpretable; the mean-intake and high-consumer figures differ by more than an order of magnitude. Marker only. 0.25
Artificial colors (red 40, caramel color) No adult outcome evidence. The pediatric behaviour signal (Southampton/McCann 2007, n=297 children) does not transfer to adults [R]. Marker only. 0.25
Titanium dioxide Colonic preneoplastic lesions (aberrant crypt foci) at 10 mg/kg bw/day in Bettini 2017. Whose exposure that is matters: Bettini justified 10 mg/kg as a child exposure estimate; adult intake is 0.2-1 mg/kg bw/day, so the study dose sits ~10-50x above an adult, not at a realistic adult dose. The same study’s 0.2 mg/kg arm produced no significant effect, and the aberrant-crypt-foci finding was not replicated (Blevins 2019, a larger rat study, found no such lesions). EFSA’s 2021 “no longer safe” verdict rested on unresolved genotoxicity rather than on the ACF result [V]. 0.25
Calcium propionate Tirosh 2019 STM, n=14: 1 g with a mixed meal acutely raised glucagon/FABP4/norepinephrine [V]. 1 g far exceeds a bread-serving dose; no chronic human data. Not detected by the scorer — it is in no detector list and scores 0. Listed here for completeness, not because it is weighted. (0 — not implemented)

Two items in the processing block (6 points, separate from the 4 above) are often mistaken for additive flags:

Ingredient Weight (of 1.0)
Processed meat (bacon, sausage, deli, “uncured”, “cultured celery”, …) 0.7
Liquid sugar (corn syrup, HFCS, cane syrup, agave nectar, juice concentrate, glucose syrup, invert sugar, molasses) 0.3
Protein isolates/concentrates/hydrolysates, collagen, gelatin 0.2

Caveat on the liquid-sugar detector. The evidence that justifies a separate liquid-sugar penalty is beverage evidence: DiMeglio 2000 compared calories delivered as a drink against the same calories as a solid and found compensation of -17% versus +118%. That finding is about the physical form in which the sugar is consumed. The detector, however, fires on syrup names in an ingredient list (corn syrup, cane syrup, agave, molasses, glucose syrup), and a syrup cooked into a teriyaki glaze is eaten as part of a solid meal. Sugar in a sauce is solid-food sugar and should be justified through the added-sugar block, not through beverage compensation data. Resolved 2026-09-06: the liquid-sugar term was removed from the processing block for solid meals and is now applied only to the 25 genuine beverage rows in the corpus, where the compensation evidence does apply. See tension 1 under “Open design tensions” at the end.

Dose multiples, stated with denominators

A multiple without a denominator cannot be checked, and the three that circulate most in this area have very different denominators depending on whose intake you use.

Additive Animal effect dose Human denominator Multiple
BHA 2% of diet, ~1,000 mg/kg bw/day EFSA 2011 adult mean 0.03-0.12 mg/kg/day ~8,000-33,000x
BHA same EFSA 2011 adult P95 0.08-1.12 mg/kg/day ~900-12,500x
Polysorbate-80 1% in drinking water, ~1,500-2,000 mg/kg bw/day EFSA 2015 toddler high consumer 24.5 mg/kg/day ~60-80x
Polysorbate-80 same water, dose estimated at ~1,500-2,500 mg/kg bw/day adult denominator: Shah et al. 2017 (FDA CFSAN), upper-bound mean intake at maximum use levels, P80 ~10-20 and CMC ~20-40 mg/kg bw/day ~75-250x (1,500/20 to 2,500/10). The EFSA 2015 toddler high-intake denominator gives ~60-80x.
Carboxymethylcellulose Chassaing 2022 human dose 15 g/day, ~214 mg/kg bw/day at 70 kg EFSA 2018 toddler P95 up to 506 mg/kg/day, and Shah 2017 adult upper bound 20-40 mg/kg bw/day ~0.4x the toddler P95, 5-11x the adult upper bound

The CMC line is the one that changes a conclusion. The “10-30x realistic intake” figure that circulates has no primary denominator on file, so it is not quoted here. Against Shah 2017’s adult upper bound of 20-40 mg/kg bw/day, 15 g/day (~214 mg/kg bw/day at 70 kg) is 5-11x. Against EFSA’s 2018 high-consumer estimate of up to 506 mg/kg bw/day in toddlers it is ~0.4x (https://pmc.ncbi.nlm.nih.gov/articles/PMC7009359/). The “wide safety margin” framing depends on which consumer you pick.

Tier 3 — drop the penalty; avoiding these is not evidence-based

Ingredient Why it’s fine
Xanthan, guar, locust/carob bean gum Fermentable soluble fibers. Partially hydrolyzed guar is a therapeutic fiber. Doses in a sauce are sub-gram; in-vitro screens use irrelevant concentrations [R].
Cellulose, methylcellulose Insoluble/non-fermented fiber. Methylcellulose is Citrucel. Chassaing 2022 [V] used carboxymethylcellulose, a different molecule, at 15 g/day. The circulating “10-30x realistic intake” multiple has no published denominator and is not used. Against retrieved denominators the trial dose is 5-11x the FDA upper-bound adult estimate and ~0.4x an EFSA high-consuming toddler’s, per the dose-multiple table above.
MSG Geha 2000 multicentre DBPCT in ~100-130 self-identified sensitive people: responses inconsistent, not reproducible, only at large doses without food [V]. Glutamate is glutamate (tomatoes, parmesan). Critically: MSG is ~12% sodium vs salt’s 39%, and partial substitution cut sodium by 31-61% across the four recipes tested (Halim 2020, https://www.ncbi.nlm.nih.gov/pmc/articles/PMC7540316/). Flagging MSG actively opposes a sodium-reduction goal.
Sulfites Relevant to sulfite-sensitive asthmatics only [R]. Otherwise handled by sulfite oxidase. Zero penalty absent asthma.
Sodium benzoate, potassium sorbate No human outcome evidence at food doses [R].
Modified food/corn starch EFSA 2017: no safety concern, no ADI needed [R]. It is starch; RS4 resistant starch is also “modified starch.” Counts only as refined carbohydrate if a top-3 ingredient.
Soy protein isolate/concentrate Reed 2021 meta-analysis, 41 clinical studies (not 38): no effect on testosterone or estradiol in men [V]. Slightly lowers LDL. Note the tension: on this evidence it should carry no penalty, but score.py gives isolates +0.2 in the processing block — not as a safety flag but as a marker that protein was added to a matrix rather than coming from food. For a high-protein goal that is arguably the wrong sign, and it is the one place where the rubric and the code disagree on purpose.
“Natural flavors” A disclosure/transparency issue, not a health one. FEMA GRAS self-review is a legitimate governance concern that has produced no human outcome evidence [R]. Keep as a processing marker only.
Maltodextrin Glucose polymer, GI commonly quoted as ~85-105 (unsourced: this range could not be traced to a primary reference and should not be quoted as measured), but 2 g in a sauce alongside 60 g of rice is noise. Rodent work used 5% in drinking water [R]. Penalize only if top-5, and then as added sugar.

Seed oils — drop the metric entirely

Canola, soybean, sunflower and safflower oil get no penalty. The accurate summary is narrower than “the evidence points the other way”: there is no harm signal, biomarker cohorts lean protective, and the randomized mortality data are null.

  • Marklund 2019 Circulation: pooled 30 cohorts, n=68,659, using tissue/blood linoleic acid biomarkers rather than self-report — higher LA associated with lower total CVD and CV mortality [V].
  • Johnson & Fritsche 2012 (15 RCTs) [V] and Su 2017 (30 RCTs, n=1,377, CRP SMD 0.09, 95% CI -0.05 to 0.24, a null rather than a reduction) [V]: no rise in CRP or IL-6 with higher linoleic acid in healthy people (https://pubmed.ncbi.nlm.nih.gov/28752873/).
  • Hooper 2020 Cochrane: the 17% fewer cardiovascular events is for reducing saturated fat by any means, 12 trials, n=53,758, not for PUFA replacement specifically. The PUFA-replacement subgroup is 7 trials, n=3,895, RR 0.73, and the review found no significant difference between replacing saturated fat with PUFA versus with carbohydrate [V] (https://pubmed.ncbi.nlm.nih.gov/32827219/). The earlier “specifically with PUFA replacement” attached the large trial’s precision to a much smaller subgroup.

Where skeptics have a real point: Ramsden 2016 BMJ recovered the Minnesota Coronary Experiment (n=9,423) — cholesterol fell, mortality did not improve [V]; the Sydney Diet Heart re-analysis (n=458) showed harm [R]. Caveats: Sydney’s margarine was trans-fat-laden, MCE had short exposure and heavy dropout, and both were corn-oil monotherapy at extremes. The omega-6:omega-3 ratio hypothesis has no RCT support — absolute EPA/DHA intake is what matters [R].

Oxidation products matter for repeatedly heated frying oil, not a tablespoon of canola in a microwaved meal. Fat still counts, but through energy density and saturated fat — not oil provenance.

Optional small bonus for extra-virgin olive oil (PREDIMED n=7,447, ~30% MACE reduction, with known randomization irregularities) [R].


The scoring rubric (v2)

This section describes exactly what [meal-service comparison removed; it is a separate private evaluation] implements. The code is the ground truth; if the two disagree, the code is right and this is stale.

v1 (a nutrient block worth ~76%, a processing block worth ~10% scored by counting UPF markers, a ~14% additive block, step bands and up to +6 in bonuses) is superseded. It was retired mainly because the marker count measured disclosure depth rather than food quality: a vendor publishing a top-level component list sails through a scan that a full FDA-style statement fails, for the same food. The step bands also made a meal’s score jump discontinuously at arbitrary cut points. There are no bonuses in v2.

Weights

Relative, not shares of 100. The nutrient block sums to 110 per column and is renormalised (below); the ingredient block adds 10 more.

Component Active adult (the adult) General population
Protein 24 12
Fiber 18 20
Energy density 16 22
Sodium 12 16
Added sugar 12 12
Potassium 10 10
Saturated fat 10 10
Meal size 8 8
Nutrient block total 110 110
Processing (PROCESSING_WEIGHT) 6 6
Additives (ADDITIVE_WEIGHT) 4 4

Ramps, not bands

Every nutrient component is a linear ramp between a good endpoint (penalty 0.0) and a bad endpoint (penalty 1.0), clamped at both ends — frac(v, good, bad) in the code. No steps, no cliffs. Everything except added sugar and meal size is normalised per calorie.

Component Quantity good (0.0) bad (1.0)
Sodium mg per kcal 0.8 1.8
Protein g per 100 kcal 6.0 3.0
Fiber g per 100 kcal 1.2 0.3
Added sugar g per meal (absolute) 5 25
Saturated fat g per 100 kcal 0.8 2.0
Energy density kcal per g 1.1 2.0
Potassium mg per 100 kcal 250 80

Energy density is a monotonic ramp, and the evidence review did not recommend that. The review’s recommendation for the adult was a target band. Low energy density is good up to a point, but a monotonic “lower is always better” ramp rewards water-heavy, low-satiety meals, and for a [calorie target removed] active adult an extremely low energy density is a problem, not a win. The scorer instead implements a ramp (1.1 good, 2.0 bad) plus a separate meal-size band, which partially recovers the intent by penalising meals that come out too small. The ramp-plus-band combination approximates a band but is not one, and the two mechanisms can pull against each other. Decided 2026-09-06: keep the ramp. A true two-sided band on kcal/g would need a guessed lower bound — there is no dose-response evidence for a floor on energy density — whereas the 450-750 kcal meal-size band is anchored to the adult’s ~5.4 trays/week and already blocks the failure mode the band was meant to prevent (a dilute 260 kcal tray winning on energy density alone). Recorded as decided, not open. See “Open design tensions”.

Potassium is scored as an instrument for intact food matrix, not as a nutrient target and explicitly not as a blood-pressure play — there is no cheap potassium-dense filler the way maltodextrin is a cheap carb filler. Regressed on the rest of the panel it gives R² = 0.432, so 57% of its variance is not explained by the panel. That is not the same as 57% being useful signal: an unknown part of the residual is label noise, rounding in published potassium figures, and vendor-to-vendor reporting inconsistency. Unexplained variance sets an upper bound on the independent information, not a measurement of it.

Meal size is an adequacy band rather than a ramp: no penalty at all between 450 and 750 kcal, then frac(kcal, 450, 220) below the band and frac(kcal, 750, 1100) above it. It exists only to stop the per-calorie normalisation from ranking 260 kcal snacks above real meals. The band is deliberately gentler than the 700-950 kcal the literature review implies, because the adult eats roughly one tray a day, not three.

Added sugar is published by only 196 of 816 meals. Total sugar is used as a proxy only when a sweet-sauce indicator appears in the dish name or ingredients; otherwise the term is dropped and renormalised away.

Processing block — 6 points

Re-scoped from a generic marker count to the two signals the cohort literature actually incriminates, plus one marker of reconstituted protein. Fractions sum, then cap at 1.0.

Signal Weight
Processed meat 0.7
Protein isolate / concentrate / hydrolysate / collagen / gelatin 0.2
Liquid sugar — beverage rows only 0.3

Liquid sugar left the solid-food processing block on 2026-09-06. DiMeglio 2000’s -17% versus +118% compensation figures are about calories consumed as a drink. Syrup cooked into a sauce is eaten as part of a solid meal, so the beverage-compensation rationale does not transfer to it, and that sugar is already the added-sugar block’s job. For a solid meal the processing block is now processed meat 0.7 + protein isolate 0.2, still capped at 1.0, still worth 6 points.

The corpus does contain genuine beverages, 25 rows of smoothies, wellness shots and protein shakes [meal-service comparison removed; it is a separate private evaluation], and for those the beverage evidence does apply, so the detector still fires on them. Beverage status is matched on dish name and category, never on ingredients (a smoothie’s ingredient list looks like a fruit bowl’s), with the same word-boundary rule the rest of the scorer uses: “Sesame Shakeup Salad” and “Strawberry Milkshake Oats” are eaten with a fork and a spoon and must not be routed here. As of 2026-09-06 none of the 25 beverage rows declares a liquid-sugar term, so the carve-out currently fires on nothing; it is kept so the rule is right when a sweetened drink enters the corpus.

The syrup terms remain in SWEET_INDICATORS, where they do unrelated work: gating the total-sugar proxy when added sugar is not published. Removing them from the processing block does not weaken that gate.

Two matching rules are load-bearing. “Uncured” and “cultured celery powder” count as processed meat — the phrase is a positive predictor that nitrite is present, so a naive matcher gets the sign backwards. Vegetable nitrate sources (spinach, beet, arugula, celery juice) are the same anion with the opposite sign and are never routed to that detector.

Additive block — 4 points

Fractions sum, then cap at 1.0. See the Tier 2 table above for the evidence behind each.

Signal Weight
Erythritol / xylitol at dose 0.5
Potassium bromate 0.5
Inorganic phosphates 0.25
TBHQ 0.25
Carrageenan, sucralose, aspartame, acesulfame, caramel color, red 40, titanium dioxide, BHA, BHT (any one) 0.25

Phosphates and TBHQ were cut from 0.6 and 0.4 on 2026-09-06 (tensions 2 and 3 below). TBHQ is not in the catch-all list — it has its own row at the same 0.25, and re-adding it to OTHER_ADDITIVE restores the double-count that made it the largest effective penalty in the block.

Phosphate matching enumerates the specific inorganic salts rather than matching bare “phosphate”, which would catch modified starches (distarch phosphate) and mineral fortificants; lecithin is an organic phospholipid and must never enter that detector.

Putting it together

  1. Compute a penalty fraction 0-1 for each component the meal actually publishes.
  2. Renormalise over present fields only: nutrient_score = 100 × (1 − Σ wᵢpᵢ / Σ wᵢ), summed over present components only — never over the full 110. A service that does not publish saturated fat gets neither credit nor blame for it.
  3. Fold in the ingredient block as a span, with span = 6 + 4 = 10 and extra = 6·proc + 4·add:

    • Ingredient list present: score = (nutrient_score + (span − extra)) × 100 / (100 + span) A clean list earns the full 10 back; a dirty one gives it up.
    • No ingredient list at all: score = nutrient_score.

    The second rule is the fix for a real defect (2026-09-05). Previously a missing list set extra = 0, which was a silent free pass worth up to 10 points and flattered the 12 meals that publish no panel [meal-service comparison removed; it is a separate private evaluation]. Missing is now neutral, consistent with how missing nutrients are handled.

    Structural edge case: a disclosed penalty can still beat no list at all. Missing and present score the same exactly when extra = (100 - nutrient_score)/10, so below that point the meal with a list comes out ahead. A meal that lists processed meat can therefore score slightly above an otherwise-identical meal with no ingredient list once the nutrient score falls under roughly 58. That is a property of renormalising over present fields rather than a bug. The 10-point span enters the denominator only when a list exists, and a weak nutrient score gains more from that renormalisation than the processing penalty takes away. It affects only the 12 meals that publish no ingredient list.

  4. Clamp to 0-100.

A published 0 is a measurement, not a gap. 0 g protein, 0 mg potassium and 0 g fiber are all real values in this corpus, and each is scored at the worst end of its ramp rather than skipped. Only None means missing. Fixing this (2026-09-05) moved four 0 g protein meals down: a “digestion shot” product fell from rank 2 to rank 60 within its source [meal-service comparison removed; it is a separate private evaluation]. kcal is the divisor for six components, so a missing or zero kcal skips every per-calorie term rather than dividing by zero.

Two traps that will silently corrupt any comparison

1. Disclosure depth. Services publish ingredients at different depths. A top-level component list (“Chicken Thigh, Andouille Sausage, Chicken Base”) will sail through an additive scan that a full FDA-style statement (“Coconut Milk [Coconut Extract, Water, Citric Acid, Sodium Metabisulfite]”) fails — for the same food. Scanning both with one regex measures transparency, not quality, and rewards the most opaque vendor. Always measure what share of lists contain bracketed sub-ingredients before comparing additive rates. Below ~25%, the service cannot be scored on additives at all; mark it ?.

2. Missing fields flatter the opaque. A service publishing no saturated fat or ingredients cannot lose points on those axes. Always report field coverage alongside any score, and never let a ? be silently treated as a zero.

NOVA / ultra-processing

Better than additive-counting as a single population-level proxy, because it captures energy density, eating rate and hyperpalatability. But a poor discriminator among prepared meals — nearly every commercial ready meal is NOVA 3 or 4, and inter-rater reliability is poor even among experts. Braesco 2022, 159 raters, 120 marketed foods: Fleiss kappa 0.32-0.34, conventionally “fair”. Kappa is chance-corrected agreement and does not translate into a raw fraction of agreements between raters; that reading of it is wrong. The raw pattern is more specific: only 3 of the 120 foods were classified identically by every rater, while 90 of the 120 drew a 91% NOVA-4 assignment rate. Raters converge hard on the ultra-processed category and disagree mostly about everything else, which is exactly why NOVA fails to discriminate within a corpus of prepared meals [V] (https://pmc.ncbi.nlm.nih.gov/articles/PMC9436773/).

Use a continuous marker count instead of a 1-4 label. Do not use “ingredient count > N” — it correlates but has no basis in NOVA.

The strongest single studies worth knowing

  • Hall 2019 [V] — n=20 inpatient crossover, 14 days/arm. Ultra-processed diet matched for presented calories, sugar, fat, sodium and fiber still drove +500 kcal/day and ~1 kg swing in 2 weeks. The strongest randomized evidence that eating an ultra-processed diet raises energy intake. It is not clean evidence that processing per se is the operative variable: the arms differed in energy density and palatability as well as processing, and the NIH 4-arm factorial that tried to isolate processing put the processing-only contribution at +128 kcal/day with a confidence interval spanning zero (see headline 1). Hall 2019 establishes the effect; it does not establish the mechanism.
  • Dicken 2025 [V] — n=55, 8 weeks/arm, free-living, both arms guideline-compliant. Minimally processed produced roughly double the weight loss. Note published critiques in Nature Medicine on effect size [V].
  • Reynolds 2019 Lancet [V] — 185 prospective studies + 58 RCTs (https://pubmed.ncbi.nlm.nih.gov/30638909/). The two halves measure different things and should not be quoted as one body of evidence. The 15-30% lower all-cause and cardiovascular mortality, highest vs lowest fiber intake, is from the prospective cohorts, at GRADE moderate certainty. The 58 RCTs (n=4,635) measured body weight, blood pressure and total cholesterol only, and no randomized mortality endpoint exists here. The dose-response statement that greater intake brings greater benefit is reported for cardiovascular disease, type 2 diabetes, colorectal cancer and breast cancer. 25-29 g/day is a floor, not an optimum: the risk curve keeps falling above it, and Reynolds 2019 says so explicitly. Still the strongest diet-outcome evidence here, but it is cohort evidence carrying the mortality claim.
  • Morton 2018 BJSM [R] — 49 RCTs, n≈1,863. Lean-mass benefit plateaus around 1.6 g/kg/day protein.
  • Sodium, honest magnitude. The effect is 1 to 6 mmHg systolic, depending on baseline diet and the size of the sodium reduction, not a flat one to two. DASH-Sodium (n=412): high-to-low sodium on the control diet lowered systolic 6.7 mmHg overall and ~5.6 mmHg in normotensives; the same sodium change on the DASH diet was worth ~3.0 mmHg, because an already-good diet leaves less for sodium reduction to fix (https://pubmed.ncbi.nlm.nih.gov/11136953/). Graudal 2020 gives -1.14 mmHg in normotensive whites, the low end of the range and the origin of the understated shorthand this document used to carry. Filippini 2021 gives -2.3 mmHg per 100 mmol/day, linear with no threshold (https://pmc.ncbi.nlm.nih.gov/articles/PMC8094404/). Outcome data: TOHP follow-up found cardiovascular events RR 0.75 (0.57-0.99) and all-cause mortality RR 0.80 (0.51-1.26, not significant) (https://pmc.ncbi.nlm.nih.gov/articles/PMC1857760/); SSaSS found death RR 0.88 (0.82-0.95) with a potassium-enriched salt substitute, which is a combined sodium-and-potassium intervention rather than sodium reduction alone (https://pubmed.ncbi.nlm.nih.gov/34459569/). Potassium supplementation is reported to have no statistically significant effect in normotensives against ~5.3 mmHg in hypertensives; the ~5.3 mmHg hypertensive figure stands, but the “3 trials, n=757” attribution is unverified; that pairing could not be confirmed against a source and should not be quoted. Sodium is worth weighting moderately, not maximally, and athletes lose sodium in sweat. It still discriminates usefully because prepared meals run 800-1,200 mg/serving.

Open design tensions

Places where the adversarial review found the rubric and the scorer disagree, or where a weight is not supported by the evidence behind it. Recorded so the next change to the scorer starts from a known list rather than rediscovering them. Tensions 1-4 were settled on 2026-09-06; 5 and 6 are still open.

# Tension Resolution
1 Liquid-sugar detector fires on solids. The beverage-compensation evidence (DiMeglio 2000) is about drinks; syrup in a sauce is solid-food sugar. APPLIED 2026-09-06. The 0.3 liquid-sugar term was removed from the processing block; solid-meal PROCESSING is now processed meat 0.7 + protein isolate 0.2, cap 1.0, weight 6 unchanged. Detection is retained for the 25 beverage rows (smoothies, wellness shots, protein shakes; [meal-service comparison removed; it is a separate private evaluation]), matched on name and category. Syrup terms stay in SWEET_INDICATORS so the added-sugar fallback gate is untouched. Effect: 41 of 816 meals lost a processing penalty; processing-zero share 88.8% → 93.2%.
2 Inorganic phosphates weighted 0.6, the largest in the additive block, on evidence that weakened on re-checking: Gutiérrez 2015 moved nothing but two bone-turnover markers, the MR finding is null except for valvular disease, and the EFSA ADI exceedance is in children, not adults. APPLIED 2026-09-06. Weight 0.6 → 0.25. The enumerated compound list and PHOSPHATE_EXCLUDE are unchanged. Effect: 29 of 816 meals (all 29 phosphate hits) carry a smaller additive penalty; additive-zero share unchanged at 94.4%, mean additive penalty 0.030 → 0.018.
3 TBHQ effectively weighted 0.65 through a double-count in both the TBHQ detector and OTHER_ADDITIVE, making it the largest single additive penalty in practice, on one lab’s mouse allergy model with no published NOAEL/ADI margin. APPLIED 2026-09-06. Weight 0.4 → 0.25 and TBHQ removed from OTHER_ADDITIVE, so a hit contributes 0.25 exactly once (verified against a synthetic meal whose ingredient list is only “TBHQ”). No meal in the current corpus declares TBHQ, so no score moved — this is a correctness fix for future data.
4 Energy density is a monotonic ramp where the evidence review recommended a target band for the adult; the scorer approximates a band with ramp-plus-meal-size instead. DECIDED 2026-09-06: keep the ramp. Energy density stays a linear ramp (1.1 → 2.0 kcal/g) and the 450-750 kcal meal-size band remains the partial compensation. A true two-sided band would require a guessed lower bound with no dose-response behind it; the kcal band is anchored to the adult’s ~5.4 trays/week and already blocks the failure mode (a dilute low-calorie tray winning on density alone). Recorded as decided, not open.
5 Potassium bromate weighted 0.5 as Tier 2 though it has no human data at all; it is a precautionary flag on a free-to-avoid ingredient, sitting in a table of human-evidence rows. Relabel as precautionary; keep or drop the weight deliberately, not by table position.
6 Erythritol 0.5 vs carrageenan 0.25 inverts study quality. Erythritol rests on a non-randomized n=10/arm trial; carrageenan rests on a randomized double-blind crossover (n=20, 14 days). Re-rank so weight tracks design quality, or document why dose realism outranks it.
  • /research/health/composition/ — what separates two meals with identical nutrition panels
  • /research/health/patterns/ — ultra-processed food epidemiology and diet-quality scoring
  • The scoring implementation and the meal-service comparison it was built for are not published. [meal-service comparison removed; it is a separate private evaluation]