Working research document, copied from the author’s notes on 2026-09-06. Rough, long, and unedited apart from removing personal details. The finding page summarizes it.

Theories and anomalies in the food-processing evidence

A skeptical-scientist pass over the anomalies in /research/health/rubric/. The goal is not to defend the rubric. It is to find the theories that explain the whole pattern of evidence including the parts that do not fit, to check whether each theory already exists in the literature, and to name the cheapest decisive test for each.

Built 2026-09-06. Every claim below carries an epistemic label. I did not run any experiment; the only things I “measured” in this session were public records (a clinical-trial registry API, journal abstracts).

Epistemic labels used throughout

Label Meaning
[VERIFIED-SESSION] I pulled the number or text from a primary source in this session and it is quoted below.
[VERIFIED-ABSTRACT] Taken from the paper’s own abstract, retrieved this session. Full text not read.
[SECONDARY] From a search summary or a secondary description. Treat as a lead.
[HYPOTHESIS] My inference. Not measured.
[SPECULATION] Plausible mechanism with no direct evidence either way.

0. Three corrections to the rubric, before anything else

These come first because the rubric is wrong or overstated on three points, and two of the eight “anomalies” partly dissolve once they are fixed.

0.1 Hall 2019 did NOT match energy density in the sense that matters

[VERIFIED-SESSION] Three different energy-density numbers get quoted about Hall 2019 and an earlier version of this section conflated them. Corrected:

Quantity Ultra-processed Unprocessed Matched?
Total presented energy density, kcal/g 1.024 1.028 yes, by design
Total consumed energy density, kcal/g 1.36 1.09 no, p < 0.0001
Non-beverage energy density, kcal/g 1.957 1.057 no, about 85% higher

The design matched what was offered, including beverages. What subjects actually ate differed by 25% in energy density, and the solid-food fraction differed by 85%. The matching was achieved partly by using beverages as vehicles for dissolved fiber supplements in the ultra-processed arm. Hall and colleagues said so themselves: “Because beverages have limited ability to affect satiety, the ~85% higher energy density of the non-beverage foods in the ultra-processed versus unprocessed diets likely contributed to the observed excess energy intake.”

So the framing “Hall 2019 found +508 kcal with total energy density matched” is true only of presented totals and is substantively misleading. The variable that the 2026 factorial identifies as dominant is precisely the variable Hall 2019 left unmatched.

But do not read that as “processing is only a proxy for energy density.” Non-beverage energy density in a real food supply is not an exposure that floats free of processing; it is largely produced by it, through drying, fat and sugar addition, water removal, extrusion, and matrix disruption. In causal terms, a factorial that holds energy density fixed across processing arms estimates a controlled direct effect of processing net of its own mediator. It answers “what is left of processing once you block the pathway it mostly works through”, not “was processing ever causal”. Correspondingly, when a cohort’s UPF hazard ratio attenuates on adding energy density, that is a demonstration of mediation, not of confounding, unless energy density is argued to be prior to processing, which it is not. Monteiro and colleagues’ 2025 Lancet series paper already names energy density, hyperpalatability, soft texture and matrix disruption as the mechanisms through which ultra-processing is proposed to act, so “it works through energy density” is the mainstream causal claim rather than a refutation of it. https://www.thelancet.com/journals/lancet/article/PIIS0140-6736(25)01565-X/abstract

Where this leaves the proxy question. Energy density is the dominant proximal mediator of processing’s effect on short-term intake. This does not adjudicate whether processing itself is causal, only through what it acts, and the intervention implication is the same under both readings: lowering non-beverage energy density is the lever, whether processing is described as “the underlying cause, acting through ED” or as “a correlate whose only real content is ED”.

0.2 The 2026 factorial trial numbers, verified from the registry

[VERIFIED-SESSION] NCT05290064, “Effect of Ultra-processed Versus Unprocessed Diets on Energy Metabolism”, NIH Clinical Center. Randomized crossover, n = 38 actual, four 1-week test diets, completed 2025-08-15, results first posted 2026-08-12. Registry-posted only; I found no peer-reviewed publication of these results.

The manipulated variable is stated in the registry as non-beverage energy density, plus hyperpalatable-food content, plus NOVA processing.

Least-squares mean energy intake (SE 108.5 kcal/day for every arm):

Diet kcal/day
UPF, high energy density, high hyperpalatable (HH) 3424.3
UPF, high energy density, low hyperpalatable (HL) 3265.9
UPF, low energy density, low hyperpalatable (LL) 2604.4
Unprocessed, low, low (UNF LL) 2476.0

Pre-specified contrasts:

Contrast Isolates Estimate (95% CI) p
UPF HH vs UNF LL total effect +948.4 (811.9 to 1084.8) <0.0001
UPF HL vs UPF LL non-beverage energy density +661.6 (524.9 to 798.3) <0.0001
UPF HH vs UPF HL hyperpalatability +158.4 (21.9 to 294.9) 0.023
UPF LL vs UNF LL processing per se +128.4 (-8.2 to 265.0) 0.065

The rubric’s numbers are right. Two secondary outcomes are more interesting than the primary and the rubric does not mention them:

Meal eating rate, grams per minute: least-squares means 39.21 (HH), 39.02 (HL), 39.04 (LL), 38.43 (UNF LL), SE 2.52 for every arm. Pairwise differences run from 0.19 to 0.78 g/min with confidence intervals of roughly plus or minus 2.76. All reported p-values are 0.99, and the fact that they are identical to two decimal places across contrasts of different magnitude suggests a multiplicity adjustment rather than four independent tests. These are registry-posted figures, not peer-reviewed ones.

Read that correctly: the interval is wide enough to contain eating-rate differences that would matter, so this is a precision-limited null, not a demonstration that eating rate is fixed. And critically, texture was not a manipulated factor in this trial. The four diets varied in processing, energy density and hyperpalatable-nutrient content, not in hardness or oral-processing demand. Flat g/min across arms that were never designed to differ in texture is consistent with the oral-processing model, not a test of it.

Meal palatability, 0 to 100 VAS: 65.24 (HH), 62.41 (HL), 66.67 (LL), 64.18 (UNF LL). All contrasts p >= 0.6. The low-hyperpalatable ultra-processed diet scored numerically highest.

Two consequences:

  1. The texture / oral-processing-speed mechanism (Forde and colleagues) is not what carried the effect in this trial, because the trial did not manipulate texture. Grams per minute did not move, and nothing in the design predicted it would. Whatever hyperpalatability did to add 158 kcal, the trial gives no evidence it did it by changing chewing demand, and no evidence it did not.
  2. “Hyperpalatability” as operationalised here is not behaving as a hedonic construct. Subjects did not rate the hyperpalatable diets as more pleasant. Any causal story that runs through “it tastes better so I eat more” is in tension with the trial’s own palatability data. [HYPOTHESIS] that the arms were built on Fazzino’s nutrient-cluster definition specifically: the registry names hyperpalatable-food content as a manipulated factor but I did not verify from the registry record which operationalisation was used, and the inference that it is the fat-sodium / fat-sugar / carbohydrate-sodium cluster definition is mine.

0.3 Rubric headline 4 is correct; CHD does not go fully to null

[Withdrawn on review: an earlier version of this section claimed the rubric’s “1.11 goes to 1.00” was probably a conflation of the total-UPF estimate with an after-exclusion estimate. That was wrong. The full text of Mendoza et al. 2024 reports both numbers, and the rubric quoted them correctly. The error was mine, made from the abstract alone after the publisher returned 403 on the full text.]

[VERIFIED-SESSION] Mendoza et al. 2024, Lancet Regional Health Americas (three US cohorts: NHS, NHSII, HPFS; n = 206,957; 16,800 CVD cases), full text at https://pmc.ncbi.nlm.nih.gov/articles/PMC11403639/

Exposure CVD CHD Stroke
Total UPF, Q5 vs Q1 1.11 (1.06-1.16) 1.16 (1.09-1.24) 1.04 (0.96-1.12)
UPF with SSB and processed meat removed 1.00 (0.96-1.05) 1.06 (1.00-1.13) not the headline

So the rubric’s headline 4 stands as written for CVD: remove sugary drinks and processed meat from the UPF variable and the composite cardiovascular association goes from 1.11 to 1.00. The one refinement worth carrying is that CHD does not go fully to null: it drops from 1.16 to 1.06 (1.00-1.13), a lower bound sitting exactly on 1. Something beyond SSB and processed meat is still contributing to coronary risk, at roughly a third of the original effect size. Stroke was never significant either way.

A separate and additional observation. The same paper’s subgroup analysis points in several directions at once:

UPF subgroup CVD HR
Sugar- and artificially-sweetened drinks 1.18 (1.09-1.29)
Processed meat 1.13 (1.04-1.23)
Savoury snacks 0.91 (0.85-0.98)
Bread and cold cereals 0.94 (0.88-1.00)
Yoghurt and dairy desserts 0.90 (0.82-0.99)

Savoury snacks (crisps, crackers, extruded snacks) inversely associated with CVD, with a confidence interval excluding 1. That is a real anomaly and it is a different one from the headline-4 point, not a substitute for it. It is taken up in Anomaly 8 and Hypothesis H5, where the sign inconsistency across cohorts and outcomes gets its own treatment.


Anomaly 1. UPF cohort associations survive nutrient adjustment, but ward trials say it is almost all energy density

Candidate explanations

A. The adjustment is not doing what it looks like it is doing. Adjusting for “diet quality” (HEI-2015, AHEI) and for a handful of nutrients does not adjust for energy density, eating rate, portion architecture, or meal timing. HEI is built out of food-group servings and a few nutrients; two diets with identical HEI can differ 2-fold in kcal/g. [HYPOTHESIS] The cohort literature has mostly not adjusted for the variable the ward trials say is causal, so “survives adjustment for diet quality” is not evidence against the energy-density explanation. This is the single most important point in this section and I could not find a cohort that adjusts UPF associations for dietary energy density as a continuous covariate.

Important qualifier on A, and it applies everywhere energy density appears in this document. If the UPF hazard ratio attenuates when energy density is added to the model, that is not evidence that the UPF association was spurious. Energy density of the food supply is downstream of processing, not a competing cause of it, so adjusting for it is adjusting for a mediator. The correct reading of attenuation is “processing acts through energy density”, which is the causal model Monteiro and colleagues state in the 2025 Lancet series, not “processing was a proxy all along”. The re-analysis below is still worth running, but its output is a mediation estimate and should be reported as one. Explanation A’s real force is narrower than it first appears: it says the cohort literature has not decomposed the pathway, not that the pathway is illusory.

B. Residual confounding by socioeconomic position and health behaviour. UPF intake correlates with smoking, income, education, physical activity, and alcohol. Adjustment is imperfect by construction because these are measured with error while UPF is measured with error in a correlated way.

C. Reverse causation. People with early disease change diet. Cohorts handle this with lag analyses, inconsistently.

D. Something real in the matrix that neither nutrients nor the ward trials capture. The ward trials measure energy intake over 1 to 2 weeks. They cannot measure 20-year cancer or CVD incidence. A matrix effect that does nothing to short-term appetite but does something to, say, colonic epithelium over decades would be invisible in Hall-style designs and visible in cohorts. This is a real logical gap and the ward trials do not close it.

E. Measurement-error asymmetry. FFQs measure UPF with different error structure than they measure nutrients. If UPF is measured better than diet quality (it is arguably easier to recall “I drink 2 sodas a day” than to reconstruct fiber grams), then adjusting for the noisier variable leaves the better-measured one carrying the signal. This is a well-known artifact and it predicts exactly the observed pattern.

What the evidence says

  • [SECONDARY] Multiple cohorts report attenuation-but-not-elimination after HEI adjustment: NHANES 2003-2018, 10-point higher %kcal from UPF associated with 9% higher all-cause mortality, reduced but still significant after diet quality adjustment. https://pubmed.ncbi.nlm.nih.gov/39608567/
  • [SECONDARY] UPF and mortality in type 2 diabetes “independent of diet quality”. https://ajcn.nutrition.org/article/S0002-9165(23)66025-3/fulltext
  • [VERIFIED-SESSION] The ward-trial decomposition (section 0.2) puts processing per se at +128 kcal/day, p = 0.065, CI crossing zero. Note the CI: the trial cannot exclude a 265 kcal/day processing effect. “Not significant at n = 38” is not “shown to be zero”. The rubric’s phrasing (“Processing is real as a warning sign, not as a cause”) is stronger than the data.
  • [SECONDARY] Lancet 2025 UPF series paper 1 (Monteiro and colleagues) tests three hypotheses: displacement of established diets, deterioration of diet quality, and increased chronic disease risk, and concludes all three are supported. Note that hypothesis 1 (displacement) is not a claim that UPF is intrinsically toxic; it is compatible with explanation A above. https://www.thelancet.com/journals/lancet/article/PIIS0140-6736(25)01565-X/abstract

Cheapest decisive test

Re-analysis, no new data collection. In NHANES, UK Biobank, or the Nurses’ cohorts, refit the UPF-mortality model adding dietary energy density (kcal/g, computed from the same 24h recalls, beverages excluded) as a continuous covariate, and report the attenuation as a mediation analysis, not as a confounding check: energy density is a consequence of processing, so the quantity being estimated is the indirect effect through energy density and the controlled direct effect net of it. Use a formal mediation estimator rather than a naive coefficient comparison. Energy density is computable from existing recall data with zero new measurement. If the UPF HR collapses toward 1 the way the intake effect collapsed in the ward trial, anomaly 1 is explained mechanistically and the intervention target is energy density. If it does not, explanation D gets much more interesting.

Second test: negative-control outcomes. It has been done, and it is uncomfortable. [VERIFIED-ABSTRACT] Morales-Berstein and colleagues 2024 (Eur J Nutr, EPIC) used accidental death as a negative-control outcome for UPF and found UPF positively associated with it. https://pmc.ncbi.nlm.nih.gov/articles/PMC10899298/

That is close to a direct measurement of the residual-confounding floor, and it comes out nonzero. An exposure cannot cause accidental death through any plausible dietary mechanism, so the association there is confounding, and by extension some fraction of the disease-outcome associations in the same cohort is confounding too. This is the strongest single piece of support for explanation B in this document. It does not tell you what fraction of the CVD or mortality HR is confounded, because the confounding structure need not be identical across outcomes, but it converts “there might be residual confounding” from a rhetorical move into a measured quantity. [Withdrawn on review: an earlier version of this section said “I could not find this done for UPF.” It has been done.]

The remaining free extension is to widen the negative-control panel (trauma mortality, hip fracture from falls in the young) and to report negative-control HRs alongside primary HRs as standard practice in UPF cohort papers.


Anomaly 2. Hall 2019 versus the 2026 factorial

Reconciled in principle; the coefficient-stability check awaits publication of the 2026 diet compositions and body-weight data. See section 0.1 and 0.2. Reconciliation, in one line: Hall 2019 matched total presented energy density including beverages; the 2026 factorial manipulated non-beverage energy density; the non-beverage energy density in Hall 2019 differed by 85% and was never matched. The two trials are consistent on that reading, but “consistent” is not “solved” until the 2026 diets’ kcal/g values are published and the implied intake-per-unit-energy-density coefficient can be compared across the two trials. Hall said this in the 2019 paper’s own discussion and it was the explicit motivation for designing NCT05290064.

Residual puzzle worth keeping. Hall 2019 produced +508 kcal/day; the 2026 non-beverage energy density contrast alone produced +662, and the full contrast +948. If energy density is the whole story, why is the 2019 effect smaller than the 2026 energy-density-only effect, when 2019 also had processing and hyperpalatability differences stacked on? [HYPOTHESIS] Because 2019’s ultra-processed arm was matched on presented fiber, sugar, fat, sodium and macros, which constrains how extreme the energy-density contrast can be, whereas the 2026 HL and LL diets were designed to maximise the energy-density contrast within a single processing category. Different contrast magnitudes, same coefficient. This is checkable the moment the 2026 diet compositions are published: if the HL-to-LL kcal/g gap is larger than the 2019 gap in proportion to the intake difference, the coefficient is stable and the story holds.

Second residual puzzle, and it is a good one. A post-hoc reanalysis of Hall 2019 (Brunstrom and colleagues, “human nutritional intelligence”) reports that subjects consumed substantially more food mass on the unprocessed diet while consuming fewer calories, with low-energy-dense components under 1.0 kcal/g making up about half of unprocessed meal mass, and claims a model combining macronutrient “blend index” and micronutrient-seeking explains about 88% of the energy difference. [SECONDARY], units not verified by me. https://pmc.ncbi.nlm.nih.gov/articles/PMC12975374/

That matters because it is a rival explanation to pure energy density that makes a different prediction: under nutritional intelligence, people are seeking micronutrients and stop when they get them, so fortifying an energy-dense ultra-processed diet to unprocessed-equivalent micronutrient density should reduce intake. Under pure energy density, fortification should do nothing. That is a clean, cheap, decisive ward trial.

Cheapest decisive test

Ward trial, 2 arms, 1 week each, n = 20 crossover. Both arms ultra-processed, both matched on non-beverage energy density and hyperpalatability. Arm A at typical UPF micronutrient density. Arm B fortified to match an unprocessed diet’s micronutrient-per-calorie profile (potassium, magnesium, folate, vitamin C, carotenoids). Outcome: ad libitum kcal/day. Nutritional intelligence predicts B < A by a few hundred kcal. Energy density predicts B = A. This is the same infrastructure Hall already has, and it is the single most informative follow-on experiment I can identify in this whole document.


Anomaly 3. Fiber’s dose-response is the strongest in nutrition, and fiber supplements do not reproduce it

Candidate explanations

A. Fiber is a marker of intact plant food, not the active agent. The nutrient travels with cell walls, potassium, magnesium, polyphenols, resistant starch, lower energy density, slower eating, and with the kind of person who eats beans.

B. Isolated fiber is a chemically different exposure. Psyllium, inulin, methylcellulose and wheat bran are not “fiber”; they are five substances with different viscosity, fermentability, and gel-forming behaviour. Pooling them and finding null is a category error.

C. Trials were the wrong duration and wrong outcome. Adenoma recurrence over 3 to 4 years is not 25-year mortality. Absence of mortality RCTs is absence of evidence.

D. Dose. Cohort top-versus-bottom fiber contrasts are roughly 15 to 35 g/day from food. Supplement trials add 7 to 15 g on top of a background diet.

What the evidence says

  • [SECONDARY] Reynolds 2019 Lancet: 185 prospective studies plus 58 RCTs, 15-30% lower all-cause and CV mortality, dose-response, benefit continues above 25-29 g/day. Cited in the rubric. Note: Reynolds 2019’s RCT arm excluded fiber-supplement-only trials, so its 58 RCTs are not a test of whether isolated fiber reproduces the food-fiber effect; this meta-analysis cannot be read as evidence either for or against explanation A on that specific question, only on the food-fiber dose-response itself.
  • [SECONDARY] DART (Burr 1989, n = 2,033 post-MI men, randomized advice design): the cereal fibre advice arm had slightly HIGHER mortality, not significant. This is the closest thing to a hard-outcome fiber RCT and it does not go the cohort’s way. https://pubmed.ncbi.nlm.nih.gov/2571009/
  • [SECONDARY] Polyp Prevention Trial: low-fat, high-fiber, high-fruit-and- vegetable diet, no reduction in adenoma recurrence, and still null at 8-year continued follow-up. https://aacrjournals.org/cebp/article/16/9/1745/176913/
  • [SECONDARY] Wheat Bran Fiber trial: no difference in recurrent adenomas with high-fiber supplement. https://aacrjournals.org/cebp/article/11/9/906/166726/
  • [SECONDARY] Structure beats composition in the acute setting: matched whole-grain and refined wheat milled products showed no difference in glycemic response or gastric emptying, while intact kernels, thick oats (>0.6 mm) and unmilled rice do attenuate glycemia. Milling to flour abolishes the whole-grain advantage. https://ajcn.nutrition.org/article/S0002-9165(22)00218-0/fulltext , https://jn.nutrition.org/article/S0022-3166(22)00044-X/fulltext , https://diabetesjournals.org/care/article/43/2/476/36090/

That last bullet is the strongest single piece of evidence for explanation A, and it is a randomized crossover, not a cohort. Same grams of fiber, same nutrients, different physical structure, different physiology. And when structure is equalised by milling, whole grain stops being special.

Verdict

Explanation A is well supported, B and C are also true, and they are not mutually exclusive. The honest statement is: “fiber intake from food” is a composite exposure of which the isolated polysaccharide is one minor component. The rubric’s claim that fiber “has the strongest outcome evidence of anything on this list” is defensible only if “fiber” is read as “fiber-rich intact plant food”. As a target for a scoring function operating on nutrition labels, grams of fiber is a fine proxy. As a claim about a molecule, it is not supported.

Cheapest decisive test

Existing-data re-analysis first. In any cohort recording both dietary fiber and supplement use (NHANES, UK Biobank, NIH-AARP), compare the mortality dose-response for food fiber versus supplemental fiber, per gram, in the same model. If the food-fiber slope is steep and the supplement slope is flat at equal grams, explanation A is strongly supported and it costs nothing. If the supplement slope matches, the molecule hypothesis survives. I found no analysis that has done this head-to-head per gram, which is surprising.

Then the feasible RCT. Not a mortality trial. A 12-week randomized 3-arm crossover: (i) 40 g/day fiber from intact legumes and whole kernels, (ii) 40 g/day fiber from the same foods milled to flour and reconstituted to identical macro and micronutrient composition, (iii) 40 g/day from mixed supplements. Outcomes: ad libitum energy intake, postprandial glucose AUC, fecal SCFA, LDL, blood pressure. Arm (ii) is the crucial one and nobody has run it at this scale. If (i) beats (ii) with identical composition, structure is the active ingredient and the field’s central nutrient is a proxy.


Anomaly 4. Potassium as an unfakeable marker of intact plant matrix

Is this in the literature?

Partly, and not in the form the rubric uses it. What exists:

  • [VERIFIED-ABSTRACT] Urinary potassium is used as a recovery biomarker for potassium intake and, more loosely, as a marker of vegetable intake. The correlations are from 24-hour urine collections, not spot urine: Spearman 0.50 in men and 0.48 in women for vegetables, and 0.25 in men / 0.35 in women for fruit. https://pmc.ncbi.nlm.nih.gov/articles/PMC10857367/ [Withdrawn on review: an earlier version cited “spot-urine correlations 0.37 to 0.46 with usual vegetable intake” against this paper. Both numbers came from a different source and neither is a spot-urine vegetable correlation.]
  • [VERIFIED-ABSTRACT] The 0.37 and 0.46 figures belong to Freedman et al. 2015: 0.37 is the FFQ-to-biomarker correlation for potassium and 0.46 is the corresponding figure for potassium density from 24-hour recalls. Same paper also supports validation correlations around 0.49 for the sodium:potassium ratio. https://pubmed.ncbi.nlm.nih.gov/25787264/
  • [SECONDARY] Na:K ratio is standard in nutrient profiling and in the hypertension literature.
  • Reviews of dietary biomarkers of ultra-processed food exist (Sci Direct S2667268525000658) but I could not retrieve full text (403) and my search budget ran out before I could confirm whether potassium appears there.

What I could not find in the time available: any paper that treats potassium as an instrument for intact food matrix on the grounds that it is economically infeasible to fortify. [HYPOTHESIS, novelty UNVERIFIED] That specific economic framing may be original to this rubric, but the claim is a did-not-find, not a searched-exhaustively, and the underlying empirical observation is emphatically not novel: potassium density falling as the ultra-processed share of the diet rises is a standard nutrient-profile finding, usually cited to Martinez Steele and colleagues’ NHANES nutrient-profile work. Treat “original to this rubric” as an unverified claim about framing only.

Why the framing is nonetheless sound, and where it breaks

The argument is an economic one, not a nutritional one: potassium is bulky (potassium chloride tastes bitter and metallic at the doses required, potassium gluconate and citrate are expensive per mg of K), it has no functional role in texture or shelf life, and it is not a marketing claim consumers pay for. So it does not get added, whereas protein, fiber, and vitamins all do get added because each has a label claim attached. That makes potassium a rare non-gameable column. The rubric’s own regression backs this up: R-squared 0.432 against the rest of the panel, so 57% of its variance is orthogonal information.

Where it breaks, and this is a real threat: potassium chloride is the leading sodium-replacement salt and its use is growing fast under sodium-reduction pressure. A processed meal that swaps 30% of its NaCl for KCl gains potassium and the rubric rewards it twice (higher K, lower Na) for what is a single reformulation with modest independent health value. [HYPOTHESIS] This makes potassium a decaying instrument: it was unfakeable in 2010 and is becoming fakeable by 2030. Also, potassium phosphate and tripotassium phosphate are common additives in processed meat and are potassium sources, which means potassium can in principle be positively correlated with the worst category in the corpus.

Cheapest decisive test

Label-data analysis on the existing 816-meal corpus, zero new data. Split the potassium column by whether the ingredient list contains a potassium salt (potassium chloride, potassium citrate, potassium lactate, potassium phosphate, potassium sorbate). Compute potassium density in each group. If the potassium-salt-containing meals have higher K density than the intact-produce meals, the instrument is already compromised in this corpus and the weight should be conditioned on the absence of potassium additives. This is a 20-line change to score.py and it is directly actionable.

Population-level test: in NHANES, model all-cause mortality on potassium intake stratified by whether potassium came from produce or from a fortified or salt-substituted source (identifiable via FNDDS food codes). If the survival slope differs by source at equal milligrams, potassium is a marker, not an agent, which is what the rubric already assumes.


Anomaly 5. Processed meat robust, red meat contested

Candidate explanations

A. Nitrite plus heme. The mechanistically favoured story: nitrite plus heme iron generates nitrosyl-heme and N-nitroso compounds in the colon.

B. Nitrite alone.

C. Sodium. 50 g of processed meat carries roughly 622 mg sodium per the rubric. The mortality/CVD signal could ride sodium, not nitrite.

D. Heterocyclic amines and polycyclic aromatic hydrocarbons. These come from high-temperature cooking, which applies to red meat too, so they cannot by themselves explain the processed-versus-unprocessed difference.

E. The eater, not the meat. Processed meat intake is a strong marker of smoking, low income, low vegetable intake, and low physical activity across every Western cohort. Unprocessed red meat is much less socially patterned. This alone predicts a robust processed-meat signal and a contested red-meat signal, with no chemistry involved.

F. Measurement. Processed meat is easy to recall and count in discrete units (2 slices of bacon, 1 hot dog). Unprocessed red meat is reported as portions of a continuous quantity with far worse recall accuracy. Differential measurement error attenuates the red-meat coefficient toward null and leaves processed meat’s intact. [HYPOTHESIS] This is my favoured partial explanation and I have seen it argued informally but did not find a formal quantification.

What the evidence says

  • [VERIFIED-ABSTRACT] Animal work is messier than an earlier version of this bullet claimed. In the npj Science of Food 2023 rat study of nitrite reduction and removal in cured meat, preneoplastic lesion burden tracked neither total fecal N-nitroso compounds nor fecal nitrosyl iron cleanly. The more consistent correlate of lesion promotion was luminal lipid peroxidation. https://www.nature.com/articles/s41538-023-00228-9 , https://www.ncbi.nlm.nih.gov/pmc/articles/PMC6893523/ [Withdrawn on review: the earlier text said promotion “did not track total N-nitroso compounds but did track nitrosyl-heme specifically”. The second half is not supported by the paper’s own abstract.] This matters for the verdict below: if lipid peroxidation is the proximal driver, the mechanistic story for processed meat shifts away from the nitrite-plus-heme chemistry and toward heme iron as a pro-oxidant in the colonic lumen, which red meat also delivers, which in turn weakens the chemical explanation for the processed-versus-unprocessed split and correspondingly strengthens explanations E and F.
  • [SECONDARY] Heme-induced biomarkers of colon-cancer promotion were not modulated by nitrite intake in one rat study, which argues against nitrite being the switch. https://pubmed.ncbi.nlm.nih.gov/23441609/
  • [VERIFIED-ABSTRACT] NutriNet-Sante separates additive nitrite from processed meat and still finds signals: nitrite additives associated with cancer (Chazelas 2022 IJE, https://academic.oup.com/ije/article/51/4/1106/6550543) and with type 2 diabetes (Srour 2023 PLOS Medicine, https://journals.plos.org/plosmedicine/article?id=10.1371/journal.pmed.1004149). But the reported correlation between additive-nitrite intake and processed-meat-food intake is about 0.73, which is far too high to claim separation. At r = 0.73 the two coefficients cannot be independently identified in a Cox model without strong assumptions.
  • [VERIFIED-ABSTRACT] The same NutriNet-Sante group’s 2026 preservative paper (n = 108,723, 1,131 incident type 2 diabetes cases) tested 17 additives, found 13 associated before correction and 12 after FDR correction (ascorbic acid, E300, dropped once corrected), including citric acid, HR 1.29 (1.11-1.51); sodium ascorbate (vitamin C), HR 1.41 (1.19-1.66); alpha-tocopherol (vitamin E), HR 1.30 (1.05-1.60); and rosemary extracts, HR 1.20 (1.02-1.41). These associations persisted after adjustment for total %UPF, inter-additive correlations were reported as low, and the paper estimates preservatives mediate 17% of the overall UPF-T2D association. https://pmc.ncbi.nlm.nih.gov/articles/PMC12780003/ [Corrected on review: an earlier version gave only “12 preservatives” with no case count, correction detail or persistence-after-%UPF-adjustment finding.] A separate UK Biobank study of 37 additive markers found colourings (HR 1.24) and flavours (HR 1.20) associated with outcomes while emulsifiers were not flagged. https://pmc.ncbi.nlm.nih.gov/articles/PMC12572789/ Vitamin E, vitamin C, citric acid and rosemary extract are not credible independent diabetogens at food-additive doses. Their appearance, and their persistence after %UPF adjustment, is the fingerprint of a shared latent exposure that %UPF adjustment does not fully remove. See Anomaly 7 and Hypothesis H4.
  • [SECONDARY] NutriRECS (Johnston 2019, Annals) rated the red and processed meat evidence low-certainty under GRADE and recommended no change; the methodological dispute is about whether GRADE is the right instrument for nutritional epidemiology, not about the data. https://pubmed.ncbi.nlm.nih.gov/31569235/

Verdict

[HYPOTHESIS] Best current theory: the processed-meat signal is a superposition of a modest real chemical effect (heme plus nitrite, and sodium) on top of a substantial confounding-plus-measurement effect (E and F), and the red-meat signal is the same modest chemical effect without the confounding boost and with worse measurement, so it fails to reach significance consistently. This explains the whole pattern including IARC’s classification split (Group 1 for processed, Group 2A for red) without requiring processed meat to be qualitatively different chemistry.

Cheapest decisive test

The natural experiment is already running, though weaker than “legislated” implies. [Corrected on review: an earlier version said France “legislated stepwise reduction” of nitrite.] France issued a 2023 government action plan, not binding legislation: 20-30% cuts in nitrite levels, elimination of nitrites from chipolata sausages specifically, and a 5-year timeline, following the ANSES 2022 report. Major producers (Herta, Fleury Michon) launched nitrite-free hams voluntarily around the same period, while the rest of the EU did not follow suit. https://www.food-safety.com/articles/7541-france-to-gradually-reduce-nitrate-in-cured-meats , https://www.foodwatch.org/en/banning-added-nitrites-and-nitrates-in-food

A difference-in-differences on French versus matched-EU colorectal cancer incidence, keyed to product-level nitrite reformulation dates and national consumption data, is the closest thing to a randomized nitrite withdrawal that will ever exist. Latency is the problem: colorectal cancer needs 10 to 20 years, so the readout is 2035 or later. The interim readout is much cheaper: fecal N-nitroso compound and nitrosyl-heme excretion in French versus matched non-French consumers of equal processed-meat quantity. There is already a method for it (https://www.ncbi.nlm.nih.gov/pmc/articles/PMC11311235/). n = 200 per group, a few months, and it directly separates “nitrite” from “processed meat”.

Cheaper still, and available today: re-run any large cohort’s processed-meat model with a negative-control exposure design. If processed meat associates just as strongly with an outcome it cannot cause, the confounding floor is measured directly.


Anomaly 6. Ultra-processed bread and cereal associate with LOWER diabetes

The finding, verified

[VERIFIED-ABSTRACT] EPIC, n = 311,892, 10.9 years, 14,236 type 2 diabetes cases. Each 10% increment of intake from UPF: HR 1.17 (1.14-1.19). But by subgroup, breads, biscuits and breakfast cereals (HR 0.65), sweets and desserts, and plant-based alternatives (HR 0.46, 0.26-0.82) were associated with LOWER incidence, and the same paper reports savoury snacks going the other way, HR 2.77 (1.09-7.05) per 10% of intake, and animal-based products HR 2.25 (1.96-2.57). Substituting UPF with minimally processed food or with processed food was associated with lower incidence. https://pmc.ncbi.nlm.nih.gov/articles/PMC11551512/

The bread and plant-based-alternative point estimates (0.65 and 0.46) are large enough that a genuine protective effect of that size, from swapping in commercial bread or a plant-based meat alternative, is not biologically plausible. Implausibly large inverse associations sitting next to an implausibly large positive one (savoury snacks at 2.77, in the same model) argue for artifact operating in both directions at once, not for one subgroup being real and the rest being noise.

Note the part the rubric omits: sweets and desserts are also inversely associated. That is much harder to explain than bread.

[VERIFIED-ABSTRACT] And in the three US cohorts, savoury snacks are inversely associated with CVD. (Lancet Reg Health Am 2024, abstract quoted in section 0.3.)

So it is not one odd category. Bread, breakfast cereal, sweets, desserts, yoghurt, and crisps all point the wrong way in large, well-conducted cohorts.

Candidate explanations

A. Substitution / compositional artifact, does not apply as stated. [Withdrawn on review: an earlier version of this explanation said the reference category for a higher bread share was “soda and processed meat”. That is wrong for how these models are actually specified.] Both the EPIC paper and the three-US-cohort paper enter all UPF subgroups (SSB, processed meat, bread and cereals, sweets and desserts, savoury snacks, yoghurt and dairy desserts, and so on) simultaneously in the same regression, each with its own coefficient. Under that specification the implicit comparator for “a higher share of intake from bread” is not soda or processed meat, which already have their own terms in the model; it is everything not separately accounted for as a UPF subgroup, which in practice is dominated by minimally processed and unprocessed food. So “bread is protective” cannot mean “bread is less bad than soda”; it means, at face value, “more bread and less minimally-processed food is associated with lower risk”, which is a harder result to wave away, not an easier one. This sharpens the anomaly rather than dissolving it.

B. Fortification. Industrial bread and cereal are the main vehicles for folic acid, iron, B vitamins, and in many countries fiber. Fortified UPF genuinely carries micronutrients that home baking does not.

C. Fiber. Commercial wholemeal bread and bran cereals are among the largest fiber sources in Western diets, and they sit in NOVA 4 because of emulsifiers and dough conditioners.

D. Reverse causation, especially for sweets. People with rising HbA1c and prediabetic symptoms cut sweets first. Over a 10-year cohort this manufactures a protective association for sweets. [HYPOTHESIS] This is my leading explanation for the sweets-and-desserts result specifically, and it is testable by lag analysis.

E. NOVA misclassification. Braesco 2022 found Fleiss kappa 0.32-0.34 among expert raters, and only 3 of 120 marketed foods classified identically by all raters. Bread is the single most contested category in NOVA (supermarket sliced bread is NOVA 4, artisan bread is NOVA 3, and coders disagree constantly). Random misclassification biases toward null; systematic misclassification can flip a sign.

F. Residual confounding in the healthy direction. Breakfast-cereal eaters eat breakfast, and breakfast-eating is a robust marker of regular schedules and health behaviour.

Cheapest decisive test

Three re-analyses on EPIC or equivalent, all free:

  1. Lag analysis. Exclude the first 4, 6, 8 years of follow-up and re-fit the sweets-and-desserts subgroup. If the inverse association weakens toward null as the lag grows, reverse causation (D) is doing the work.
  2. Absolute-intake instead of share-of-intake models. Refit subgroups in grams/day with total energy adjusted by the residual method rather than as a share of total intake. If the bread inversion disappears, explanation A is confirmed as an artifact of the compositional denominator.
  3. Fortification stratification. EPIC spans countries with and without mandatory folic acid and iron fortification. If the bread inversion is present only in fortifying countries, explanation B carries it. This is a ready-made natural experiment inside a dataset that already exists.

Anomaly 7. Emulsifier cohort signals exist, trials are null or mixed

The evidence, verified

[VERIFIED-ABSTRACT] Cohort (NutriNet-Sante, Lancet Diab Endo 2024, n = 104,139, 1,056 cases, mean 6.8y): total carrageenans HR 1.03 per 100 mg/day; tripotassium phosphate 1.15 per 500 mg/day; E472e 1.04 per 100 mg/day; sodium citrate 1.04 per 500 mg/day; guar gum 1.11 per 500 mg/day; gum arabic 1.03 per 1000 mg/day; xanthan gum 1.08 per 500 mg/day. https://www.thelancet.com/journals/landia/article/PIIS2213-8587(24)00086-X/fulltext

[VERIFIED-ABSTRACT] Trial (CGH 2026, n = 60 healthy, 10 per arm, self-described as exploratory, 2-week emulsifier-free run-in then 4 weeks of CMC, polysorbate-80, carrageenan, soy lecithin, native rice starch, or placebo in brownies): total cholesterol fell during the emulsifier-free run-in itself (p = 0.00006), before any emulsifier was reintroduced. During the 4-week treatment phase: SCFA concentrations lower with CMC (and directionally with others); carrageenan increased transcellular permeability versus baseline (p = 0.04); no differences in fecal calprotectin, CRP, serum LBP, cholesterol, other metabolic markers, or serum inflammatory and cardiometabolic proteins between placebo and emulsifiers. https://www.cghjournal.org/article/S1542-3565(25)00698-6/fulltext

[SECONDARY] ADDapt (randomized, double-blind, n = 154, Crohn’s, 8 weeks, emulsifier restriction with re-supplementation): clinical remission 49.4% versus 30.7%, p = 0.019. [Corrected on review: an earlier version recorded this as “49% versus 31%” with no p-value.] Control-arm doses: 2.5 g/d carrageenan, 3.6 g/d CMC, 0.8 g/d polysorbate-80. https://academic.oup.com/ecco-jcc/article/19/Supplement_1/i262/7967009

The decisive observation, and why “decisive” overstates the CGH trial

Look at what the cohort incriminates: guar gum and xanthan gum, the two substances the rubric places in Tier 3 as harmless fermentable fibers, carry HRs of 1.11 and 1.08. And the same group’s 2026 preservative paper incriminates vitamin E, vitamin C, citric acid and rosemary extract.

[HYPOTHESIS] A model in which each of these ~20 chemically unrelated substances has an independent diabetogenic effect is not credible. A model in which all of them are noisy indicators of a single latent variable, “grams of industrially packaged food per day”, predicts exactly this pattern: every additive gets a hazard ratio proportional to its correlation with the latent factor, and none of it survives an intervention that holds packaged-food intake constant.

But the CGH trial is underpowered for anything but a large effect, not evidence of no effect. [Corrected on review: an earlier version said the effects “vanished” under the CGH intervention and called ADDapt “irrelevant”. Neither framing survives looking at the sample sizes.] Ten people per arm for four weeks, in a trial the authors themselves describe as exploratory, can detect only large, consistent effects; a true effect the size of the cohort hazard ratios (roughly 10-15% relative risk per exposure increment) is nowhere near detectable at that n. The run-in period is the tell: total cholesterol fell significantly (p = 0.00006) merely from removing emulsifiers from the background diet for two weeks, before any emulsifier was reintroduced, which shows the trial can register a real physiological shift when the exposure contrast is large enough. The null result during the 4-week treatment phase is consistent with an underpowered design testing a smaller, more realistic exposure contrast, not with “emulsifiers have no effect”.

ADDapt is the stronger evidence here, not the weaker one. It is randomized, double-blind, n = 154 (an order of magnitude larger than any CGH arm), and found a clinical remission difference of 49.4% versus 30.7% (p = 0.019) from emulsifier restriction. Its population is Crohn’s rather than healthy adults, and it did not hold total food matrix constant the way the CGH trial did, so it cannot be generalised to a healthy adult without qualification. But it should not be weighed as weaker evidence than a 10-per-arm exploratory null just because its population is further from the general population: the honest synthesis is that emulsifier restriction has a demonstrated effect in a susceptible population and an untested (not a demonstrated null) effect in healthy adults.

Cheapest decisive test

Zero-cost test on published data: for each of the ~20 incriminated additives, plot its published hazard ratio (per SD of intake) against its correlation with total ultra-processed food grams in the same cohort. Under the latent-factor model these fall on a line through the origin, and the residual scatter is the substance-specific effect. Under substance-specific toxicity there is no such relationship, and biologically inert substances (vitamin E, citric acid, rosemary extract) sit at HR = 1. NutriNet-Sante has published enough to do most of this without new data access. This single scatter plot would settle the additive question for the field.

Positive-control / negative-control additive design: pre-register a set of additives with no plausible mechanism (colour-only, at trace dose) as negative controls, and report their HRs alongside the hypothesised ones in every future additive-cohort paper. If the negative controls light up as brightly, the paper has measured the packaging, not the chemical.

The RCT that would matter and has not been run: 12 weeks, n = 150, three arms, all on an identical minimally processed background diet, with emulsifiers added at 95th-percentile population dose versus placebo, stratified by a susceptibility marker (baseline mucosal barrier function or a CARD15/NOD2-type genotype). The existing null trial is 4 weeks in 60 healthy people with 10 per arm, which cannot detect a subgroup effect.


Anomaly 8. Why sugary drinks and processed meat, and not chips or candy

The direction is inconsistent across cohorts and outcomes, not just sharper than stated

[Corrected on review: an earlier version framed this anomaly as a single consistent “chips and candy look protective” finding. It is not consistent across cohorts or outcomes, which is a different and messier problem.] [VERIFIED-ABSTRACT] In the three US cohorts, savoury snacks were inversely associated with CVD, HR 0.91 (0.85-0.98). In EPIC, sweets and desserts were inversely associated with type 2 diabetes, but the same EPIC paper finds savoury snacks running the other way for type 2 diabetes, HR 2.77 (1.09-7.05) per 10% of intake, comparable to animal-based products at 2.25 (1.96-2.57) (see Anomaly 6). https://pmc.ncbi.nlm.nih.gov/articles/PMC11551512/ So savoury snacks are protective for CVD in one cohort and among the largest positive associations for type 2 diabetes in another. [VERIFIED-ABSTRACT] Fang and colleagues 2024 add a third pattern for mortality: packaged savoury snacks, HR 1.01 (0.99-1.04), essentially null, while sweet snacks and desserts run positive in at least one cause-specific model, HR 1.21 (1.13-1.30). https://pmc.ncbi.nlm.nih.gov/articles/PMC11077436/

So the question is not “why don’t chips and candy show harm”; it is “why does the same food category flip sign across cohorts and outcomes, in a range that includes point estimates (bread 0.65, plant-based alternatives 0.46) too large to be a genuine protective effect of that food”. That pattern argues for artifact operating in both directions, not for chips-and-candy having a real, sign-consistent protective effect.

Candidate explanations, in order of my credence

A. Liquid calories are not compensated. This is the best-evidenced distinction and it is not about processing at all. DiMeglio and Mattes 2000: 450 kcal/day as jelly beans was compensated (total intake and weight flat); 450 kcal/day as soda was not (intake and weight rose). Chips and candy are solid; soda is not. https://www.nature.com/articles/0801229 (Int J Obes 24:794-800). [Corrected on review: an earlier version attached the Wiley DOI 10.1046/j.1467-789X.2003.00112.x to this citation. That DOI belongs to Almiron-Roig 2003, a different paper, cited correctly elsewhere in this document.] The rubric records the magnitude as -17% compensation for liquid versus +118% for solid; the +118% figure is [UNVERIFIED] against the primary source in this session and should be checked before being repeated as a hard number.

B. Fructose dose delivered fast to the liver. SSB delivers a large fructose bolus with no fiber and no chewing delay, driving de novo lipogenesis, uric acid production, and visceral/ectopic fat. Candy delivers similar sugar but slower and in smaller per-occasion doses. https://www.sciencedirect.com/science/article/pii/S0735109715049074 , https://pmc.ncbi.nlm.nih.gov/articles/PMC4850171/

C. Frequency and dose asymmetry. A daily soda habit is 2 to 3 servings/day for years. A daily crisps habit is 1 small bag. The exposure contrast between top and bottom quintile is far larger for SSB than for savoury snacks, so SSB has more statistical room to show an effect.

D. Confounding runs in opposite directions by social class. In US and European cohorts, SSB intake is strongly inversely socially patterned (poorer, younger, more smoking). Crisps, crackers and desserts are much more evenly distributed and in some cohorts skew toward higher socioeconomic position, which would produce a spurious protective association. [HYPOTHESIS] This is likely the main driver of the inverse associations specifically.

E. Processed meat is special for a different reason than SSB. SSB harm is metabolic and dose-driven; processed meat harm is carcinogenic plus sodium plus confounding. They land in the same UPF bucket for classification reasons, not because they share a mechanism. That two mechanistically unrelated categories carry the CVD signal is itself an argument that “ultra-processed” is not a mechanism.

Cheapest decisive test

Re-analysis: in any cohort, model CVD against percent of daily energy consumed in liquid form (all beverages except water, computed from existing recall data) and against NOVA-4 percent, in the same model, and compare the information criteria and the C-statistic improvement. This directly tests Hypothesis H2 below and requires no new data. If liquid-calorie fraction dominates NOVA-4 fraction, the classification system should arguably be replaced for cardiometabolic purposes.

Natural experiment already available: SSB taxes. Mexico (2014), Berkeley (2015), Philadelphia (2017), UK Soft Drinks Industry Levy (2018), and the UK levy in particular caused reformulation rather than only price change, which is a near-ideal instrument. Interrupted time series on dental caries and, with longer lag, on incident type 2 diabetes, keyed to levy dates, tests whether reducing liquid sugar specifically moves outcomes. Some of this is published; a harmonised multi-jurisdiction analysis is the missing piece.


Unifying hypotheses

Seven candidates. Each states the claim, whether it exists in the literature, what supports and contradicts it now, one falsifying test, and a confidence label.

Confidence scale: Established (would bet 90%+), Probable (70-90%), Plausible (40-70%), Speculative (15-40%), Weak (under 15%).


H1. Non-beverage energy density is the master variable for intake, and almost everything else in the UPF literature is a correlate of it

Claim. For short-to-medium-term ad libitum energy intake in adults, kcal per gram of the solid food presented explains most of the variance attributed to “processing”, “palatability”, “texture” and “additives”. Everything else is a proxy that loses its coefficient when energy density is entered.

In the literature? Yes, substantially. This is Rolls’ volumetrics research programme from the 1990s onward, and it is Hall’s stated interpretation of the adult’s own trials. Credit: Barbara Rolls (energy density and satiation), Kevin Hall (NCT03407053, NCT05290064). The specific framing “non-beverage” energy density is Hall’s, from the 2019 discussion section and the 2026 trial design.

Supports. [VERIFIED-SESSION] The 2026 factorial: +661.6 kcal/day from non-beverage energy density alone (p<0.0001), versus +158.4 from hyperpalatability (p=0.023) and +128.4 from processing (p=0.065). [VERIFIED-SESSION] Hall 2019’s unmatched non-beverage energy density (1.957 versus 1.057 kcal/g). The rubric’s own note that matching energy density collapses the intake SMD from 0.71 to 0.02.

Contradicts. (i) The processing contrast in the 2026 trial has a CI upper bound of +265 kcal/day, so a real 250 kcal/day processing effect is not excluded at n = 38. (ii) Energy density explains intake, and intake explains obesity, but it does not obviously explain colorectal cancer or the processed-meat signal. H1 is a theory of one outcome, not of the whole field. (iii) The Brunstrom reanalysis offers a rival account of the same data via micronutrient-seeking.

Falsifying test. Run the 2026 factorial’s UPF LL versus UNF LL contrast at n = 150 (enough to detect 130 kcal/day at 80% power). If processing per se comes in at +130 kcal/day with a CI excluding zero, H1’s “almost everything” claim is falsified and processing has an independent intake effect of about 14% the size of the total.

Confidence. [Corrected on review: an earlier version gave a single confidence rating for H1. The claim has a narrow and a strong reading and they warrant different numbers.] Narrow form (non-beverage energy density is the single biggest lever on short-term ad libitum intake, among the levers this document examines): 70%. Strong form (processing per se, additives and hyperpalatability reduce to “only a proxy” for energy density with no independent causal content): under 35%, because energy density is chiefly a mediator of processing’s effect rather than a confound that displaces it (section 0.1); the intervention implication is the same either way, since reducing non-beverage energy density is the actionable lever regardless of whether processing is described as “the cause, acting through ED” or as “a correlate whose only real content is ED”. Speculative as an explanation of chronic-disease endpoints not mediated by intake or adiposity.


H2. The fraction of calories consumed in liquid or semi-liquid form predicts cardiometabolic outcomes better than NOVA class

Claim. Replace “percent of energy from NOVA-4” with “percent of energy from liquid and semi-liquid vehicles (beverages, drinkable yoghurts, smoothies, soups, purees, sauces consumed with a spoon)” and you get a better-fitting, more mechanistically interpretable, and far more reliably coded predictor.

In the literature? The component is well established (Mattes and DiMeglio on liquid-calorie non-compensation; the entire SSB literature). Credit for the component: Richard Mattes, DiMeglio and Mattes 2000. [Corrected on review: an earlier version called the comparative framing novel. It is not fully novel.] Cordova et al. 2023 (EPIC multimorbidity; https://pmc.ncbi.nlm.nih.gov/articles/PMC10730313/) reports the same subgroup-heterogeneity pattern this framing predicts: SSB HR 1.09, animal-based products HR 1.09, breads/cereals HR 0.97. And Mendoza et al.’s SSB-and- processed-meat removal (section 0.3) is functionally a test of the same idea in reverse: strip out the liquid-sugar-heavy category and watch the composite collapse. What I did not find published is the specific nested-model C-statistic comparison proposed below; that comparison, not the general idea that liquid sugar carries a disproportionate share of the signal, is the actual increment this document adds.

Supports. [SECONDARY] DiMeglio and Mattes: matched 450 kcal/day as soda produced no compensation and weight gain; as jelly beans, full compensation and no weight gain. [VERIFIED-ABSTRACT] SSB is one of only two UPF subgroups positively associated with CVD in three US cohorts, while solid UPF subgroups (savoury snacks, bread, yoghurt) are inversely associated. [VERIFIED-SESSION] Hall 2019 achieved its energy-density “match” only by exploiting beverages, which is an unintentional demonstration that beverages behave differently from solids. [SECONDARY] NOVA’s inter-rater kappa is 0.32-0.34 (Braesco 2022) whereas “is it a liquid” has near-perfect reliability, so the proposed variable is strictly better measured.

Contradicts. (i) Processed meat is solid and carries half the CVD signal, so H2 cannot be the whole story; at best it explains the SSB half. (ii) Milk, 100% fruit juice and soup are liquids with heterogeneous or protective associations, so “liquid” alone is too coarse; the variable probably needs to be “liquid AND energy-dense AND low-protein”. (iii) [SECONDARY] Almiron-Roig 2003 reviewed the liquid-satiety literature and found it inconclusive, with some studies showing solids less satiating than liquids. https://onlinelibrary.wiley.com/doi/abs/10.1046/j.1467-789X.2003.00112.x

Falsifying test. In UK Biobank or NHANES, fit three nested Cox models for CVD and for incident type 2 diabetes: (a) covariates only, (b) covariates + NOVA-4 percent, (c) covariates + liquid-energy percent. Both (b) and (c) must also adjust for total sugar intake; without that, a liquid-energy-fraction term could simply be re-capturing total sugar dose rather than the liquid-versus- solid mechanism, and the comparison would not isolate what H2 claims to isolate. Compare change in C-statistic and AIC. If NOVA-4 percent adds predictive information after liquid-energy percent is in the model, and liquid-energy percent does not add information after NOVA-4, H2 is falsified. Cost: one analyst, existing data.

Confidence: Plausible. 45% on liquid-energy fraction beating NOVA-4 for type 2 diabetes and 35% for total CVD (because of processed meat), revised down from an earlier 55%/40% once Cordova’s and Mendoza’s results are counted as partial evidence already spent rather than fully independent future confirmation.


H3. Fiber’s benefit is structural, not chemical: the active exposure is intact plant cell walls, and grams of fiber is a proxy for them

Claim. The mortality dose-response for fiber is a dose-response for intact plant cell structure. Fiber grams index it because cell walls are made of fiber. Isolated fiber added back to a disrupted matrix restores the number on the label and little else.

In the literature? Yes, and it has a name. Credit: Anthony Fardet’s “food matrix effect” paradigm (Journal of Food Science 2024, https://ift.onlinelibrary.wiley.com/doi/10.1111/1750-3841.17139), together with Fardet and Rock’s parallel food-matrix synthesis and structure-focused reviews by Reynolds and colleagues (2020) and Kim and colleagues (2022), and the plant-cell-wall encapsulation literature in cereal science. Fardet states the holist version explicitly: chronic disease risk tracks “degradation and artificialization of food matrices” more than composition.

Supports. [SECONDARY] Milled whole wheat gives the same glycemic response and gastric emptying as milled refined wheat at matched composition, while intact kernels, thick oats and unmilled rice attenuate glycemia. This is a randomized crossover demonstration that structure, not composition, carries the effect. [SECONDARY] DART’s fibre-advice arm had non-significantly higher mortality. [SECONDARY] Polyp Prevention Trial and Wheat Bran Fiber trial both null for adenoma recurrence.

Contradicts. (i) Psyllium, an isolated fiber, does lower LDL and improve glycemia reproducibly, so isolated fibers are not inert; the mechanism is viscosity, which is a physical property, so this arguably supports a physical rather than chemical reading anyway. (ii) The hard-outcome fiber RCTs are mostly about adenoma recurrence over 3 to 4 years, an outcome and duration that may simply be wrong for detecting a mortality effect. Null there is weak evidence. (iii) Confounding: fiber intake is one of the most socially patterned dietary variables in existence, so the cohort dose-response has a large confounding component regardless of which mechanism is right.

Falsifying test. The milled-versus-intact 12-week crossover described in Anomaly 3, arm (i) versus arm (ii): identical foods, identical composition, identical fiber grams, one milled to flour and reconstituted. If ad libitum energy intake, postprandial glucose AUC and fecal SCFA are indistinguishable between intact and milled at matched composition, H3 is falsified and the molecule story wins.

Confidence: 75% that structure contributes substantially to fiber’s observed benefit (credit: Fardet, Fardet and Rock, Reynolds and colleagues 2020, Kim and colleagues 2022, and the milled-versus-intact glycemia RCTs above). 45% that it dominates over composition, i.e. that isolated fiber at matched dose is close to inert; psyllium’s reproducible LDL and glycemic effects (a physical, viscosity-driven mechanism, but still an isolated-fiber effect) cap the “dominates” reading below 50%.


H4. Additive cohort associations are a single latent factor (packaged-food grams), not N independent chemical effects

Claim. The published additive-outcome hazard ratios are generated by one underlying exposure. Each additive’s HR is approximately proportional to its correlation with total packaged-food intake, with negligible substance-specific residual. The apparent breadth of the additive literature is not N independent findings; it is one finding reported N times.

In the literature? The criticism is common in commentary (confounding by overall UPF intake is raised in every review). The specific quantitative prediction, that HRs should be a linear function of correlation-with-UPF with near-zero residual and that biologically inert additives should show indistinguishable HRs, I did not find stated or tested. Credit for the general critique: widely held. The testable form appears to be novel here.

Supports. [VERIFIED-ABSTRACT] The 2026 NutriNet-Sante preservative paper (n = 108,723, 1,131 cases) finds 12 of 17 tested preservatives associated with type 2 diabetes after FDR correction, including alpha-tocopherol (vitamin E, HR 1.30), sodium ascorbate (vitamin C, HR 1.41), citric acid (HR 1.29) and rosemary extracts (HR 1.20), all persisting after adjustment for total %UPF, with low reported inter-additive correlation and an estimated 17% mediation of the UPF-T2D association by preservatives specifically. If the method returns vitamin E and rosemary extract as diabetogens net of %UPF, the method is measuring something %UPF adjustment does not capture, not just measuring chemistry badly. https://pmc.ncbi.nlm.nih.gov/articles/PMC12780003/ [VERIFIED-ABSTRACT] The 2024 emulsifier paper incriminates guar gum (HR 1.11) and xanthan gum (HR 1.08), both of which are fermentable soluble fibers with therapeutic uses. [VERIFIED-ABSTRACT] The CGH emulsifier RCT held the background diet emulsifier-free in all arms and delivered emulsifiers through identical brownies, thereby breaking the latent factor, and found no inflammatory or metabolic effect distinguishable from a large exploratory null (see Anomaly 7 for why that trial, 10 per arm and self-described as exploratory, is underpowered rather than decisive).

Contradicts. (i) The CGH trial did find lower SCFA with CMC and increased transcellular permeability with carrageenan versus baseline. Those are real, substance-specific, dose-realistic effects, so the latent factor is not the whole story. (ii) ADDapt found clinical benefit (49.4% versus 30.7%, p = 0.019, randomized double-blind) from emulsifier restriction in active Crohn’s, which is hard to explain as pure confounding since it was randomized. H4 must therefore be qualified: it applies to healthy-adult population cohorts, not to susceptible subgroups. (iii) A UK Biobank 37-additive-marker study flags a different subset (colourings, flavours) than NutriNet-Sante flags, with emulsifiers not significant there; if a single latent “packaged-food grams” factor generated every additive-outcome association purely through correlation strength, the flagged subset should look more similar across food supplies and cohorts than it does.

H4’s central vulnerability. [Corrected on review: an earlier version treated “persists after %UPF adjustment” as settling the question in H4’s favour.] That finding only supports H4 if %UPF is a good proxy for “grams of industrially packaged food”. If %UPF is instead a poor proxy for the latent factor H4 proposes, measured with substantial error or capturing a somewhat different construct, then residual association after adjusting for %UPF is exactly what a real, largely-%UPF-independent exposure would also produce, and H4 has not been tested against that alternative. NOVA’s own kappa of 0.32-0.34 (Anomaly 6) means %UPF itself is a noisy measurement of processing, so adjusting for it will not fully remove a shared latent exposure even if H4 is true. This weakens the inference from “persists after %UPF adjustment” to “measuring something other than packaging” more than the 2026 paper credits.

Falsifying test. The scatter plot in Anomaly 7: published HR per SD versus correlation with total UPF grams, across all ~20 incriminated additives in NutriNet-Sante. H4 predicts (asserted here, not derived from any published model) an R-squared above roughly 0.7 on that line, and predicts inert substances (vitamin E, citric acid, rosemary extract) sit on the line rather than at HR = 1. If the substances with known mechanisms (CMC, P80, carrageenan) sit clearly above the line while the inert ones sit at 1, H4 is falsified and substance-specific toxicity is real.

Confidence: Plausible, roughly 50%, revised down from Probable, for the healthy-adult cohort literature, because the persistence-after-%UPF-adjustment evidence assumes %UPF is a good proxy for the latent factor, an assumption H4 has not itself tested. Weak as a claim about Crohn’s disease.


H5. Processing harm at the population level is carried almost entirely by two to three product classes, and “ultra-processed” as a category has near-zero incremental predictive value once they are entered

Claim. Sugar-sweetened beverages, processed meat, and (probably) alcohol account for essentially all of the outcome signal attributed to ultra-processing. A model containing grams/day of those three plus energy density plus fiber will fit as well as or better than any NOVA-based model, in every cohort, for every cardiometabolic outcome.

In the literature? The heterogeneity finding is now mainstream and well-cited: EPIC, the three US cohorts, and the Lancet Reg Health Americas paper all report divergent subgroup associations and explicitly recommend “deconstructing the classification”. Credit: Cordova et al. 2024 (EPIC, Lancet Reg Health Eur), Mendoza/Chen et al. 2024 (Lancet Reg Health Am). What I did not find is anyone testing the strong form: that NOVA adds zero after the top categories are entered. That specific test appears not to have been published.

Supports. [VERIFIED-ABSTRACT] In three US cohorts, only SSB and processed meat were positively associated with CVD; bread/cold cereals, yoghurt/dairy desserts and savoury snacks were inversely associated. [VERIFIED-ABSTRACT] In EPIC, breads/biscuits/breakfast cereals, sweets and desserts, and plant-based alternatives were inversely associated with type 2 diabetes. [SECONDARY] NOVA inter-rater kappa 0.32-0.34 (Braesco 2022) is genuine reliability evidence: independent trained coders applying the same system to the same foods disagree that often. [Corrected on review: an earlier version added “7.9% versus 45.9% UPF share from the same data under different systems (Martinez-Perez 2021)” as if it were more of the same evidence.] That figure is not a reliability measurement in the same sense: four different classification systems built around different intentions define “ultra-processed” differently, so disagreement across systems is partly a validity difference in what is being measured, not solely noise in measuring one thing. The kappa figure alone supports the point about NOVA’s own reliability; the cross-system percentage gap should not be stacked on top of it as independent corroboration.

Contradicts. (i) The EPIC total-UPF HR of 1.17 per 10% increment is large and the inverse subgroups do not obviously cancel it; something is generating the aggregate signal beyond the two bad classes. (ii) The subgroup analyses are compositional (shares of intake), so inverse associations may be substitution artifacts rather than evidence that those foods are benign, which weakens the inference in both directions. (iii) Ready-to-eat meat, poultry and seafood products are the strongest mortality subgroup in at least one analysis (HR 1.13), which may be a fourth class rather than a variant of processed meat.

Falsifying test. Nested-model comparison in three independent cohorts. Model A: covariates + SSB g/day + processed meat g/day + alcohol + energy density + fiber. Model B: Model A + NOVA-4 percent. If the NOVA-4 term in Model B is significant with a hazard ratio meaningfully above 1 and improves the C-statistic, H5 is falsified. Pre-register the C-statistic improvement threshold (say 0.002) before looking.

Confidence: 65% on NOVA-4 adding less than a 0.003 C-statistic improvement over Model A for CVD. 45% for type 2 diabetes and for all-cause mortality, lower because the bread/sweets/savoury-snacks subgroup pattern in those outcome domains (Anomalies 6 and 8) is larger, sign-inconsistent across cohorts, and harder to absorb into a clean three-category model than the CVD pattern is. Note on alcohol: alcohol is not a food-processing category; NOVA does not classify beverages by ethanol content, so treating alcohol as one of H5’s “two to three product classes” alongside SSB and processed meat is a category error relative to what NOVA actually measures. H5 should be read as resting on SSB and processed meat as the UPF-specific classes; alcohol belongs in Model A only as a known cardiometabolic and carcinogenic confounder, not as part of the “ultra-processed” construct being tested.


H6. The body defends a food-mass and eating-duration setpoint, not a calorie setpoint, so energy intake is mechanically proportional to energy density

Claim. People eat a habitual mass of food at a habitual rate for a habitual duration. Grams per minute and minutes per day are the regulated quantities. Calories are the passenger. Energy density is therefore not one risk factor among many, it is the exchange rate between the regulated variable and the outcome variable.

In the literature? The mass/volume component is Rolls’ “people eat a constant weight of food” observation from the 1990s, and the rate component is Forde’s energy-intake-rate work. Credit: Barbara Rolls (constant food weight), Ciaran Forde (energy intake rate, kcal/min = g/min x kcal/g). The strong combined form (both mass and duration are defended, calories are wholly derived) I have not seen stated as a single testable claim.

Supports, with an important design caveat. [Corrected on review: an earlier version called the 2026 factorial’s flat eating rate a “direct, measured demonstration of the mechanism”. It is not, for the reason below.] [VERIFIED-SESSION] The 2026 factorial’s registry-posted secondary outcome: least-squares mean eating rate in grams per minute was 39.21, 39.02, 39.04 and 38.43 across the four diets, SE 2.52 for every arm, pairwise differences 0.19 to 0.78 g/min, CIs roughly plus or minus 2.76, and every reported p-value identical at 0.99 (suggesting a multiplicity adjustment rather than four independent tests; see section 0.2). Texture was not a manipulated factor in this trial. The four diets varied in processing, energy density and hyperpalatable-nutrient content, not in hardness or oral-processing demand, so a flat g/min here is consistent with the oral-processing model, not a test of it: it is a precision-limited null on a variable the design never manipulated.

The actual test of texture comes from the same research programme’s other trials, which do manipulate it. [VERIFIED-ABSTRACT] Teo and colleagues 2023 (n = 18, energy density equalised across conditions at 1 kcal/g so only texture varied): eating rate 85% slower on the hard-texture meal, ad libitum intake 33% lower (571 kcal), with the paper’s own conclusion that “level of processing did not affect food intake” once energy density and texture were controlled separately. https://pmc.ncbi.nlm.nih.gov/articles/PMC10469122/ [VERIFIED-ABSTRACT] Teo and colleagues 2025, J Nutr (n = 69): an energy-density-by-eating-rate interaction, p = 0.024; the slow-rate, low-energy-density meal produced 573 kcal, about 50% less intake than the fast, high-energy-density meal. https://doi.org/10.1016/j.tjnut.2025.06.006 [SECONDARY] Forde and colleagues 2026, AJCN (n = 41): 369 kcal/day difference by texture, but this trial was not matched for macronutrients (fat 22 versus 33 EN%, protein 21 versus 16 EN% across arms), so texture and macronutrient composition are confounded in that result, and it should be weighted below the two energy-density-matched Teo trials.

Contradicts, and seriously. [SECONDARY] The Brunstrom reanalysis of Hall 2019 reports subjects ate substantially more food mass on the unprocessed diet (figures around 666 g versus 424 g, units not verified by me) while eating fewer calories. If total mass consumed differs by 57% between arms, mass is plainly not defended. The strong form of H6 is therefore already contradicted for total mass. [Corrected on review: an earlier version said the rate component “survives strikingly well” and narrowed H6 to eating rate being “a stable individual trait invariant to diet composition”. That narrowing is contradicted by the Teo trials above, which is the opposite of what the earlier text claimed.] Eating rate moved 85% with texture in Lasschuijt 2023 and interacted with energy density (p = 0.024) in Teo 2025, so g/min is not composition-invariant; it is a manipulable variable that, combined with energy density, arithmetically sets kcal/min (Forde’s own kcal/min = g/min x kcal/g framing, which is prior art, not a new finding here). H6 (narrow form) should read: eating rate is one input to a multiplicative relationship with energy density, not a fixed trait the body defends regardless of what is on the plate.

Falsifying test. A within-subject crossover manipulating energy density across a wide range (0.8 to 2.5 kcal/g) with all meals eaten ad libitum and instrumented plates recording gram-by-gram intake and duration. A version of this test has effectively already been run: Lasschuijt 2023 and Teo 2025 vary energy density and texture directly and find g/min moves substantially, which counts as evidence against the invariance form of H6, not for it. H6 (narrow form) predicts g/min is flat across all conditions and that between-condition variance in kcal/day is almost fully accounted for by kcal/g x duration. If g/min varies systematically with energy density, H6 is falsified; the Teo data already suggest it does. The 2026 trial is uninformative on this specific point because it never manipulated texture.

Confidence: under 20% for eating rate as a fixed, diet-composition-invariant trait; the one trial that manipulated texture directly moved g/min by 85%, and the 2026 factorial’s flat g/min is not evidence against that because texture was never varied there. Weak for the strong mass-setpoint form (contradicted by Hall 2019’s mass data). What survives at higher confidence is the much narrower, non-novel claim that kcal/min = g/min x kcal/g arithmetically, which is Forde’s original framing, not a new discovery about invariance.


H7. “Hyperpalatability” as currently operationalised is a nutrient-composition variable, not a hedonic one, and the field has mislabelled its own construct

Claim. Fazzino’s hyperpalatable-food definition (three nutrient clusters: fat plus sodium, fat plus sugar, carbohydrate plus sodium) identifies foods that increase intake, but does so without those foods being rated as more pleasant. Whatever the effect is, it is not “reward-driven eating because it tastes better”. The name is causing the field to reason about a mechanism that is not there.

In the literature? Fazzino’s definition is well-cited and explicitly built from nutrient clusters rather than from hedonic ratings (https://onlinelibrary.wiley.com/doi/10.1002/oby.22639, 62% of 7,757 FNDDS foods meet criteria). [Corrected on review: an earlier version said the critique that the construct is not hedonic “appears to be novel here”. It is not.] Rogers 2024 (Appetite) made essentially this critique already, measuring palatability directly alongside hyperpalatable-food status and finding that no metric of hyperpalatability predicted desire to eat, which cuts against a hedonic reading as much as it cuts against a “wanting, not liking” rescue of the construct. https://doi.org/10.1016/j.appet.2024.107596 Credit for the definition: Fazzino et al. 2019. Credit for the critique: Rogers 2024, with the 2026 factorial’s null palatability VAS as an independent, later data point pointing the same way.

Supports. [VERIFIED-SESSION] In the 2026 factorial, the palatability VAS was 65.24 (high-HPF UPF), 62.41 (UPF HL), 66.67 (low-HPF UPF), 64.18 (unprocessed). All contrasts p >= 0.6, and the low-hyperpalatable arm scored numerically highest. Yet the high-HPF arm ate 158 kcal/day more. [VERIFIED-ABSTRACT] Rogers 2024 measured palatability directly alongside hyperpalatable-food status across a range of foods and found no metric of hyperpalatability predicted desire to eat, an independent, earlier demonstration of the same dissociation the 2026 factorial’s VAS data show. https://doi.org/10.1016/j.appet.2024.107596 [SECONDARY] A definitional overlap analysis exists asking how nutrient clustering, NOVA and nutrient profiling relate to measured palatability (https://www.sciencedirect.com/science/article/pii/S0195666324003994), which is the right question.

Contradicts. (i) A single-item end-of-meal VAS at n = 38 is a blunt instrument; it may lack the sensitivity to detect a hedonic difference that nonetheless drives behaviour. Absence of a rating difference is not absence of a reward difference. (ii) Implicit reward measures (progressive-ratio breakpoint, wanting-versus-liking dissociation) may show effects where explicit liking does not, which is the standard finding in incentive-salience research. So H7’s strong claim (“not hedonic”) could be wrong in an interesting way: the foods may be more wanted without being more liked.

Falsifying test. Repeat the HPF-versus-non-HPF contrast with a proper reward battery: progressive-ratio operant breakpoint, forced-choice preference, and separate wanting/liking VAS (Leeds Food Preference Questionnaire), in addition to intake. If breakpoint or wanting is elevated for HPF foods while liking is flat, H7 is refined into “wanting, not liking”. If nothing moves on any reward measure and intake still rises, the effect is not reward at all, and the name should be retired in favour of something like “fat-salt-sugar co-occurrence”.

Confidence: 65% overall. High confidence that explicit liking is not the mediator (both the 2026 factorial and Rogers 2024 point the same way); lower confidence that no reward process at all is involved, since implicit wanting/liking dissociation has not been tested directly against this construct.


Scoreboard: what would move each headline in the rubric

Rubric headline Status after this pass Single test that would settle it
1. UPF is mostly a proxy for calorie density Strengthened. Verified the factorial numbers. But add the caveat that the processing contrast CI reaches +265 kcal/day. Repeat the UPF LL vs UNF LL contrast at n = 150.
2. Processed meat is the clearest bad actor Holds, with a large confounding component unquantified. Negative-control-outcome analysis in any cohort; French nitrite natural experiment for the chemical part.
3. Liquid sugar is the clean visceral-fat finding Strengthened by the subgroup pattern in Anomaly 8. Liquid-energy-fraction versus NOVA-4 nested model comparison.
4. Removing processed meat + SSB takes CVD to null Correct. [Corrected on review: an earlier version called this “probably wrong as stated” and guessed it conflated the total-UPF estimate with an after-exclusion estimate. The full text confirms the rubric quoted both numbers correctly.] Verified from the Mendoza et al. 2024 full text: CVD goes from 1.11 (1.06-1.16) to 1.00 (0.96-1.05) on removing SSB and processed meat. CHD does not go fully to null: 1.16 (1.09-1.24) to 1.06 (1.00-1.13). Done: full text retrieved (section 0.3). Remaining: whether a similar exclusion analysis has been run for the CHD-specific residual.
5. Additives are the weakest pathway Strongly strengthened. The 2026 preservative paper incriminating vitamin E and rosemary extract is close to a reductio. The HR-versus-correlation-with-UPF scatter plot.
6. Seed oils point the other way Not examined this pass. n/a
7. Fiber has the strongest outcome evidence Weakened as a claim about the molecule; unchanged as a claim about fiber-rich food. Food-fiber versus supplement-fiber slopes per gram in the same cohort.
8. Sodium matters less than its reputation Not examined this pass. n/a
10. NOVA tells you almost nothing about a specific food Strengthened. The inverse associations for bread, sweets, yoghurt and crisps are what an unreliable classifier looks like. Nested model: does NOVA-4 add anything over the two bad classes?

Implications for score.py

Only two are concrete enough to act on, and both are cheap:

  1. Condition the potassium term on the absence of potassium additives. Anomaly 4. If a meal’s ingredient list contains potassium chloride, potassium citrate, potassium lactate or potassium phosphate, potassium is no longer an instrument for intact matrix in that meal and the term should be down-weighted or dropped. Run the corpus split first to see whether it matters in practice.
  2. Consider a liquid-calorie fraction term if the corpus ever includes beverages or drinkable items. Hypothesis H2. Currently the corpus is trays, so this is probably moot.

Nothing here justifies changing the additive block. If anything, Anomaly 7 argues for shrinking it further.

Open questions I could not resolve this session

  • [Resolved on review] Whether the Lancet Regional Health Americas paper contains an exclude-SSB-and-processed-meat analysis at all: yes, full text retrieved (section 0.3). The exclusion analysis exists and the rubric’s headline 4 is correct for CVD, with a CHD residual (1.06) the rubric does not mention.
  • Whether potassium appears in the published UPF-biomarker literature as a matrix marker (search budget exhausted before I could retrieve the review at doi S2667268525000658).
  • The exact units in the Brunstrom reanalysis of Hall 2019 (per meal versus per day), which matters for how strongly it contradicts H6.
  • Whether the 2026 factorial’s diets have published energy densities, which is needed for the coefficient-stability check in Anomaly 2.

Hardest open questions (from review)

  1. What else differed between UPF HL and UPF LL in the 2026 factorial, since non-beverage energy density was manipulated by swapping which foods made up the diet, not by adding or removing a single nutrient. “+662 kcal/day from energy density alone” is a food-bundle contrast, not a single-variable manipulation, and should be described that way whenever the number is quoted.
  2. What observation would distinguish “UPF is a proxy for energy density” from “UPF causes overeating via energy density”? Both readings predict the same attenuation pattern when energy density is added to a cohort model, and this document does not yet have a test that separates them (see the “proxy question” note in section 0.1 and H1).
  3. The negative-control-outcome analysis (accidental death, Morales-Berstein 2024, Anomaly 1) exists and came out positive. Why has it not been applied to SSB and processed meat specifically before this document calls them “the clearest bad actors”? If they carry the same confounding floor as total UPF, some fraction of their apparent harm is the same artifact.
  4. Why is bread’s inverse association (HR 0.65 in EPIC) treated as an artifact of the compositional model, while SSB’s positive association from the same kind of share-of-intake model is treated as real? Both come from the same modelling choice; a principled reason is needed to trust one sign and distrust the other, not a preference for the sign that matches the rubric’s priors.
  5. How much of the scoreboard in this document rests on NCT05290064, a single unpublished, registry-only, n = 38 dataset? Several headline numbers (the 661.6 kcal/day energy-density contrast, the flat g/min secondary outcome) have no peer-reviewed backing yet.
  6. How much measurement error in %UPF would H4 require before “persists after %UPF adjustment” stops being informative? NOVA’s own kappa (0.32-0.34) is large, and H4 has not stated the error magnitude its own falsifying test assumes away.
  7. If eating rate in grams per minute were a stable, diet-composition-invariant trait, why did Lasschuijt 2023 move it 85% with a texture manipulation, in the same research programme that produced the eating-rate framework in the first place?
  8. Hall 2019 stacked processing, energy density and hyperpalatability together and found +508 kcal/day. The 2026 factorial’s energy-density-only contrast alone found +662 kcal/day. Under a simple additive H1 story, 2019 should have produced a number at least as large as any single component from 2026, not a smaller one. This document does not resolve that.
  9. On what principle is a 10-per-arm, self-described exploratory trial’s null result treated as decisive evidence against emulsifier harm, while a randomized, double-blind, adequately powered, p = 0.019 positive trial (ADDapt) is treated as irrelevant because its population is not “healthy adults”? Both are evidence; they should not receive opposite epistemic weight when both readings happen to favour the null.
  10. Which of these seven hypotheses would change any decision made by a tray-meal scoring function that has no beverages in its input space? If the honest answer is “none of them, because H1’s beverage/non-beverage split and H2’s liquid-calorie mechanism both concern a food category the scorer never sees”, that should be stated plainly rather than left implicit.

Sources

Primary records retrieved this session:

  • NCT05290064 registry record and posted results (ClinicalTrials.gov API v2): https://clinicaltrials.gov/study/NCT05290064
  • Hall 2019, Cell Metabolism: https://www.cell.com/cell-metabolism/fulltext/S1550-4131(19)30248-7
  • Energy-density detail in Hall 2019 (secondary analysis, PMC): https://pmc.ncbi.nlm.nih.gov/articles/PMC7946062/
  • Brunstrom et al., post-hoc reanalysis of Hall 2019: https://pmc.ncbi.nlm.nih.gov/articles/PMC12975374/
  • Cordova et al. 2024, EPIC, degree of processing and type 2 diabetes, Lancet Reg Health Eur: https://www.ncbi.nlm.nih.gov/pmc/articles/PMC11551512/
  • UPF and CVD, three US cohorts, Lancet Reg Health Am 2024: https://www.thelancet.com/journals/lanam/article/PIIS2667-193X(24)00186-8/fulltext
  • Emulsifiers and type 2 diabetes, NutriNet-Sante, Lancet Diab Endo 2024: https://www.thelancet.com/journals/landia/article/PIIS2213-8587(24)00086-X/fulltext
  • Preservatives and type 2 diabetes, NutriNet-Sante, Nature Communications 2026 (full text): https://pmc.ncbi.nlm.nih.gov/articles/PMC12780003/
  • Additive markers, UK Biobank, 37-marker study: https://pmc.ncbi.nlm.nih.gov/articles/PMC12572789/
  • Cordova et al. 2023, EPIC multimorbidity: https://pmc.ncbi.nlm.nih.gov/articles/PMC10730313/
  • Fang et al. 2024, packaged food subgroups and mortality: https://pmc.ncbi.nlm.nih.gov/articles/PMC11077436/
  • Rogers 2024, hyperpalatability and measured palatability, Appetite: https://doi.org/10.1016/j.appet.2024.107596
  • Lasschuijt et al. 2023, “Speed limits”, energy-density-matched texture trial: https://pmc.ncbi.nlm.nih.gov/articles/PMC10469122/
  • Teo et al. 2025, energy-density x eating-rate interaction, J Nutr: https://doi.org/10.1016/j.tjnut.2025.06.006
  • DiMeglio and Mattes 2000, liquid versus solid carbohydrate, Int J Obes: https://www.nature.com/articles/0801229
  • Five dietary emulsifiers RCT, Clin Gastro Hepatol 2026: https://www.cghjournal.org/article/S1542-3565(25)00698-6/fulltext
  • ADDapt trial abstract, J Crohns Colitis: https://academic.oup.com/ecco-jcc/article/19/Supplement_1/i262/7967009
  • Nitrites/nitrates and cancer, NutriNet-Sante, Int J Epidemiol 2022: https://academic.oup.com/ije/article/51/4/1106/6550543
  • Nitrites/nitrates and type 2 diabetes, PLOS Medicine 2023: https://journals.plos.org/plosmedicine/article?id=10.1371/journal.pmed.1004149

Supporting literature:

  • Fardet 2024, ultra-processing as a holistic food-matrix issue, J Food Sci: https://ift.onlinelibrary.wiley.com/doi/10.1111/1750-3841.17139
  • Fazzino et al. 2019, hyper-palatable food definition, Obesity: https://onlinelibrary.wiley.com/doi/10.1002/oby.22639
  • Nutrient clustering, NOVA and palatability overlap: https://www.sciencedirect.com/science/article/pii/S0195666324003994
  • Forde and colleagues, ultra-processing or oral processing, Curr Dev Nutr: https://www.ncbi.nlm.nih.gov/pmc/articles/PMC7042610/
  • Teo et al., texture-based eating rate and energy intake, AJCN 2022: https://ajcn.nutrition.org/article/S0002-9165(22)00025-9/fulltext
  • Lasschuijt et al., “Speed limits”: https://www.ncbi.nlm.nih.gov/pmc/articles/PMC10469122/
  • Matched whole-grain vs refined milled products, AJCN 2022: https://ajcn.nutrition.org/article/S0002-9165(22)00218-0/fulltext
  • Oats processing and glycemia meta-analysis, J Nutr: https://jn.nutrition.org/article/S0022-3166(22)00044-X/fulltext
  • Wholegrain particle size and glycemia, Diabetes Care: https://diabetesjournals.org/care/article/43/2/476/36090/
  • DART, Burr 1989: https://pubmed.ncbi.nlm.nih.gov/2571009/
  • Polyp Prevention Trial continued follow-up: https://aacrjournals.org/cebp/article/16/9/1745/176913/
  • Wheat Bran Fiber trial: https://aacrjournals.org/cebp/article/11/9/906/166726/
  • Almiron-Roig 2003, liquid calories and satiety, Obesity Reviews: https://onlinelibrary.wiley.com/doi/abs/10.1046/j.1467-789X.2003.00112.x
  • Fructose, uric acid and de novo lipogenesis: https://www.sciencedirect.com/science/article/pii/S0735109715049074 , https://pmc.ncbi.nlm.nih.gov/articles/PMC4850171/
  • Nitrite in vivo evidence review: https://www.ncbi.nlm.nih.gov/pmc/articles/PMC6893523/
  • Nitrite reduction/removal in cured meat, rat model, npj Sci Food 2023: https://www.nature.com/articles/s41538-023-00228-9
  • Heme biomarkers not modulated by nitrite: https://pubmed.ncbi.nlm.nih.gov/23441609/
  • Fecal nitrosocompound excretion method: https://www.ncbi.nlm.nih.gov/pmc/articles/PMC11311235/
  • NutriRECS, Johnston 2019: https://pubmed.ncbi.nlm.nih.gov/31569235/
  • France nitrite reduction policy: https://www.food-safety.com/articles/7541-france-to-gradually-reduce-nitrate-in-cured-meats , https://www.foodwatch.org/en/banning-added-nitrites-and-nitrates-in-food
  • UPF and mortality, NHANES 2003-2018: https://pubmed.ncbi.nlm.nih.gov/39608567/
  • UPF, diet quality and mortality in type 2 diabetes, AJCN: https://ajcn.nutrition.org/article/S0002-9165(23)66025-3/fulltext
  • Lancet 2025 UPF series, paper 1: https://www.thelancet.com/journals/lancet/article/PIIS0140-6736(25)01565-X/abstract
  • Nature Medicine, concerns over conclusions in a UPF trial (critique of Dicken 2025): https://www.nature.com/articles/s41591-025-04087-7
  • Urinary biomarkers for fruit/vegetable, sodium, potassium and Na:K: https://pmc.ncbi.nlm.nih.gov/articles/PMC10857367/
  • Pooled validation of self-report against potassium/sodium recovery biomarkers: https://pubmed.ncbi.nlm.nih.gov/25787264/
  • /research/health/rubric/ - the rubric this pass interrogates
  • [meal-service comparison removed; it is a separate private evaluation] - the implementation