A systematic review and meta-analysis of randomised trials

Early versus delayed fortification of human milk in preterm infants

Search date 18 August 2026. 8 randomised trials (786 infants); 7 trials (706 infants) contribute data to quantitative synthesis.

Status of this document. This is a reproducible synthesis of published summary data conducted by a single assessor, not a registered systematic review, and it should not be cited as one. Screening and extraction were not duplicated, trial authors were not contacted for missing data, and two eligible reports were available only as abstracts. The protocol (protocol.md) was fixed before extraction and every deviation is declared in section 2.6.

Plain-language summary

Fortifier is added to breast milk to supply the extra protein, energy and minerals that preterm infants need. Units have traditionally waited until an infant is tolerating close to a full volume of milk before adding it, on the concern that fortifying early might cause feeding problems or bowel injury. The competing argument is that the first weeks are exactly when the nutritional deficit builds up, so waiting has its own cost.

We found 8 randomised trials of starting fortifier early versus later, of which 7 provided usable numbers. Early fortification was associated with regaining birth weight about a day sooner (MD -1.10 (95% CI -2.09 to -0.12) days). For every other outcome we examined, including weight gain velocity, necrotising enterocolitis, death, infection, feed intolerance, time to full feeds and length of hospital stay, the trials did not show a clear difference either way, and the confidence intervals remained wide enough to include effects that would matter clinically in both directions.

The evidence is not strong. Every trial had at least some risk-of-bias concerns, the trials are small, and certainty in the results was rated low or very low for 10 of the 11 outcomes. The practical reading is that early fortification does not appear to cause the harms that motivated waiting, but the trials to date are too small and too few to establish that confidently, and the one apparent benefit is of uncertain clinical importance.


1. Background and objective

Multi-component fortifier corrects the protein, energy and mineral deficit of unfortified human milk in preterm infants. Practice has commonly withheld it until enteral feeds approach 100 mL/kg/day, on the theory that an osmotically active supplement added to early low-volume feeds increases feed intolerance and necrotising enterocolitis (NEC). The competing position is that the first one to two postnatal weeks are when the cumulative nutrient deficit accrues, so deferral guarantees a deficit in exchange for an unproven safety benefit.

The existing Cochrane review of this question (CD013392, Thanigainathan and Abiramalatha 2020, search to August 2019) included two trials and 237 infants and concluded that the evidence was insufficient to support or refute early fortification. Several trials have appeared since. This review updates that synthesis and asks, in addition, whether effects vary with how "early" is operationalised, with gestational age, and with fortifier source.


2. Methods

Methods were fixed in protocol.md before extraction and analysis. This section summarises them; the protocol is the authoritative statement.

2.1 Eligibility

Randomised or quasi-randomised trials in preterm (< 37 weeks) or low-birth-weight infants receiving enteral human milk, in which the randomised contrast was the timing of initiation of multi-component fortification with all other nutritional management intended to be equivalent between arms. Early was defined as fortification begun at an enteral volume < 100 mL/kg/day or at < 7 days postnatal age; delayed as >= 100 mL/kg/day or >= 7 days. These thresholds follow CD013392 so the included set is comparable with it. Trials whose timing contrast was confounded by a deliberately co-randomised second intervention were excluded and are discussed narratively. Trials randomising to early fortification versus an unfortified early diet, with the control arm fortified once standard criteria were met, satisfy the timing contrast and are eligible; that framing is examined in a prespecified sensitivity analysis.

PubMed/MEDLINE via E-utilities, Europe PMC (which indexes CENTRAL records), and ClinicalTrials.gov, with no date or language restriction, plus backward citation harvest from CD013392 and three subsequent reviews, and a registry-linked companion search on the registration identifier of every included trial. Search date 18 August 2026. Strings and hit counts are recorded in the protocol; every screening decision and its reason is in screening_log.csv.

2.3 Data collection

One row per trial per outcome in long format (extracted_data.csv). For every extracted number the source column records the table, sentence or abstract it came from and source_access records whether that was full text, publisher abstract, or preprint. Where a mean was reported without an SD, the SD was reconstructed from a reported SE, 95% interval, or median and IQR (Wan et al. 2014) and flagged imputed_sd; 16 data points across 4 trials required this, and every affected outcome is re-analysed without them.

2.4 Risk of bias

Cochrane RoB 2, five domains, with a written justification per domain per trial (risk_of_bias.csv). Because fortifier changes milk appearance, blinding is absent in most trials; the consequence was judged separately for objective outcomes (mortality, NEC) and subjectively assessed ones (feed intolerance, the decision to withhold a feed).

2.5 Synthesis

Dichotomous outcomes: Mantel-Haenszel random-effects risk ratios. Continuous outcomes: inverse-variance random-effects mean differences (units are common within each outcome, so a standardised difference was not needed). Between-study variance by REML. Because the number of trials per outcome is small, the Hartung-Knapp adjusted interval is the primary inferential result throughout; conventional Wald intervals, which are anticonservative here, are reported alongside in pooled_results_classicCI.csv. Heterogeneity is reported as tau-squared, I-squared with its 95% interval, Cochran Q, and a 95% prediction interval where at least three trials contribute. Trials with zero events in both arms contribute no information to a risk ratio and are reported as dropped rather than silently omitted.

2.6 Deviations from the protocol

  1. Single-reviewer screening and extraction rather than two independent reviewers. Declared in advance (protocol section 4) and mitigated by recording a verbatim source for every number, but it remains a material limitation and no inter-rater agreement can be reported.
  2. Two trials were assessed from an abstract rather than full text: Alizadeh Taheri 2017 is paywalled and Siddiqui 2025 exists only as a preprint. The protocol anticipated this (section 11) and both are flagged in every table. Only Siddiqui 2025 yielded extractable arm-level data; Alizadeh Taheri 2017 contributes to narrative synthesis alone.
  3. DerSimonian-Laird estimates were not tabulated separately. The protocol offered them for comparability with CD013392; REML with Hartung-Knapp is reported instead, with conventional Wald intervals as the comparison, because that pair is the more informative contrast at this number of trials.
  4. The gestational-age subgroup was operationalised as < 29 weeks or < 1250 g rather than the protocol's < 28 weeks, because no included trial used a 28-week boundary; this follows the entry criteria of the trials that exist.

3. Results

3.1 Study selection

The search returned 275 unique records. 248 were excluded at title and abstract, 27 reports were assessed in full, 18 were excluded with recorded reasons, and 1 registry record (NCT05251441) is awaiting classification. 8 trials (786 infants) met eligibility criteria. Of these, 7 (706 infants) contribute data to at least one pooled estimate; Alizadeh Taheri 2017 reports no arm-level numbers, SDs or event counts and contributes to narrative synthesis only.

PRISMA 2020 flow diagram
Figure 1. PRISMA 2020 flow diagram.

The most common reason for exclusion at full text was a comparison other than fortification timing: trials of fortifier product, of individualised versus standard fortification, of fortified versus permanently unfortified milk, and of routine versus selective fortification. One trial was excluded because its timing contrast was bundled with a co-randomised feed-advancement regimen (protocol section 3.3), and one report (Tucker 2025) was excluded as a secondary analysis of an already-included trial, the unit of analysis being the trial rather than the publication.

3.2 Included trials

Characteristics of included trials
Figure 2. Characteristics of included trials.

The 8 trials were conducted in the USA (5), India (2) and Iran (1). Populations range from infants born below 29 weeks (Salas 2023, mean birth weight 795 g) and 500-1250 g (Sullivan 2010) to broader preterm and VLBW groups. Definitions of "early" divide into a feed-volume threshold (fortifier begun at 20-40 mL/kg/day, or from the first feed) and a postnatal-age threshold (day 3 to day 7), with comparators at 70-100 mL/kg/day or day 10-14 respectively. Fortifier was bovine milk-derived in most trials; Sullivan 2010 and the first two weeks of Salas 2023 used a human milk-derived product. 5 of the 7 trials in synthesis report a registration identifier.

Sullivan 2010 requires comment. It is a three-arm trial designed to compare fortifier source, and only two of its arms (human milk-derived fortifier begun at 40 versus 100 mL/kg/day) form an eligible timing contrast. That contrast is a secondary, non-prespecified comparison within the trial, which is reflected in its risk-of-bias rating and in a sensitivity analysis.

3.3 Risk of bias

RoB 2 judgements by domain
Figure 3. RoB 2 judgements by domain.

No trial was at low risk of bias overall: 4 were rated high risk and 4 raised some concerns. The dominant drivers were absence of masking in trials with subjectively assessed outcomes (domains 2 and 4) and unverifiable prespecification (domain 5). Salas 2023 is rated by outcome class: high risk for growth outcomes, where the primary outcome was ascertained in 70% of randomised infants, but low risk for safety outcomes, which use full randomised denominators. Salas 2025 documents protocol drift that narrowed the intended timing contrast, which biases its contribution toward the null. Alizadeh Taheri 2017 was assessable only from a publisher abstract, so its "some concerns" ratings record absent reportable detail rather than observed flaws.

3.4 Primary outcomes

Time to regain birth weight. 4 trials, 352 infants. Early fortification shortened the time to regain birth weight by MD -1.10 (95% CI -2.09 to -0.12) days (p = 0.0379), with no detectable heterogeneity (I-squared 0.0%, tau-squared 0.000, Q p = 0.623). This is the only outcome in the review whose interval excludes no effect. It should be read against the prespecified minimally important difference of 2 days: the point estimate is about half that, and the interval's lower bound only just reaches it. The conventional Wald interval is narrower (-1.90 to -0.31), so the significance of this result is not an artefact of the Hartung-Knapp adjustment, but the estimate is fragile in sensitivity analysis (section 3.9).

Time to regain birth weight (days)
Figure 4. Time to regain birth weight (days).

Weight gain velocity. 6 trials, 604 infants. No difference was detected: MD +0.30 (95% CI -0.99 to +1.59) g/kg/day (p = 0.58). This is the most heterogeneous outcome in the review (I-squared 65.5%, 95% CI 17.4 to 85.6%; Q p = 0.0127), and the 95% prediction interval (-2.46 to +3.05) spans effects that would be clinically meaningful in either direction against a 2 g/kg/day minimally important difference. The pooled estimate is therefore not a useful summary of what a future trial should expect.

Weight gain velocity (g/kg/day)
Figure 5. Weight gain velocity (g/kg/day).

3.5 Safety outcomes

No safety outcome showed a difference between early and delayed fortification, and every interval is wide enough to contain both meaningful benefit and meaningful harm.

NEC (Bell stage >= 2). 7 trials reported the outcome; 2 (Gupta 2025, Siddiqui 2025) recorded zero events in both arms and contribute no information to a risk ratio, leaving 5 analysable trials, 516 infants and 19 events. RR 0.96 (95% CI 0.32 to 2.90), p = 0.921. The interval spans a two-thirds reduction to a near-tripling of risk. In absolute terms, control-arm risk 39 per 1000 -> 37 per 1000 with early fortification (95% CI 12 to 113). This is the outcome that motivated the practice of delaying fortification, and the honest summary is that these trials neither confirm nor exclude the concern.

NEC, Bell stage >= 2
Figure 6. NEC, Bell stage >= 2.

All-cause mortality before discharge. 3 trials, 326 infants, 16 deaths. RR 1.25 (95% CI 0.55 to 2.86), p = 0.369.

All-cause mortality before discharge
Figure 7. All-cause mortality before discharge.

Late-onset sepsis. 5 trials reported the outcome; 2 had zero events in both arms, leaving 3 analysable trials, 347 infants and 54 episodes. RR 0.77 (95% CI 0.37 to 1.63), p = 0.277.

Late-onset sepsis
Figure 8. Late-onset sepsis.

Feed intolerance. 5 trials, 404 infants. RR 1.09 (95% CI 0.78 to 1.53), p = 0.527. The definition of feed intolerance is trialist-specific and clinician-driven in unmasked trials, which is why this outcome is downgraded for indirectness as well as risk of bias.

Feed intolerance
Figure 9. Feed intolerance.

3.6 Other outcomes

Outcome Trials Infants Effect (95% CI) p I-squared
Length gain velocity (cm/wk) 4 428 MD +0.043 (95% CI -0.051 to +0.138) 0.239 36.6%
Head circumference velocity (cm/wk) 4 428 MD +0.006 (95% CI -0.117 to +0.128) 0.891 64.2%
Postnatal growth failure / EUGR 4 370 RR 1.04 (95% CI 0.92 to 1.17) 0.389 0.0%
Time to full enteral feeds (d) 6 542 MD +0.14 (95% CI -0.25 to +0.52) 0.407 0.0%
Duration of hospital stay (d) 6 542 MD +1.64 (95% CI -2.19 to +5.46) 0.321 0.0%

Time to full enteral feeds is the one outcome rated moderate certainty, because its estimate is both precise and homogeneous: MD +0.14 (95% CI -0.25 to +0.52) days across 6 trials with I-squared 0.0%. The interval excludes a difference of even half a day in either direction, so this is a genuine null rather than an uninformative one — early fortification does not delay the attainment of full feeds.

Duration of hospital stay carries the widest interval of the continuous outcomes (MD +1.64 (95% CI -2.19 to +5.46) days) and is the outcome most sensitive to analytic choices (section 3.9).

Length gain velocity (cm/week)
Figure 10. Length gain velocity (cm/week).
Head circumference velocity (cm/week)
Figure 11. Head circumference velocity (cm/week).
Postnatal growth failure / EUGR at discharge
Figure 12. Postnatal growth failure / EUGR at discharge.
Time to full enteral feeds (days)
Figure 13. Time to full enteral feeds (days).
Duration of hospital stay (days)
Figure 14. Duration of hospital stay (days).

Narrative synthesis. Alizadeh Taheri 2017 (80 infants) reports that growth indices and complication rates did not differ between arms, without arm-level numbers, SDs or event counts. It is consistent in direction with the pooled estimates but contributes to none of them.

3.7 Heterogeneity

Outcome Trials tau-squared I-squared (95% CI) Q p 95% prediction interval
Time to regain birth weight (d) 4 0.0000 0.0% (0.0 to 84.7) 0.623 -2.39 to +0.18
Weight gain velocity (g/kg/d) 6 0.9117 65.5% (17.4 to 85.6) 0.0127 -2.46 to +3.05
Length gain velocity (cm/wk) 4 0.0014 36.6% (0.0 to 78.1) 0.193 -0.108 to +0.195
Head circumference velocity (cm/wk) 4 0.0034 64.2% (0.0 to 87.9) 0.0387 -0.213 to +0.224
Postnatal growth failure / EUGR 4 0.0000 0.0% (0.0 to 84.7) 0.939 0.75 to 1.44
Time to full enteral feeds (d) 6 0.0000 0.0% (0.0 to 74.6) 0.615 -0.32 to +0.59
Feed intolerance 5 0.0000 0.0% (0.0 to 79.2) 0.708 0.69 to 1.73
NEC (Bell stage >=2) 5 0.0000 0.0% (0.0 to 79.2) 0.57 0.26 to 3.50
All-cause mortality 3 0.0000 0.0% (0.0 to 89.6) 0.85 0.16 to 9.79
Late-onset sepsis 3 0.0000 0.0% (0.0 to 89.6) 0.614 0.27 to 2.25
Duration of hospital stay (d) 6 2.4440 0.0% (0.0 to 74.6) 0.58 -4.64 to +7.92

Two outcomes carry substantial heterogeneity: weight gain velocity (I-squared 65.5%, Q p = 0.0127) and head circumference velocity (I-squared 64.2%, Q p = 0.0387). Everything else has an I-squared point estimate at or near zero — but the 95% intervals around those estimates reach 75-90% in every case, so homogeneity is not established anywhere in this review. At four to six trials per outcome, I-squared is estimated too imprecisely to be read as evidence of consistency, and the prediction intervals are the more honest summary: for weight gain velocity a future trial could plausibly observe anything from -2.46 to +3.05 g/kg/day.

Note that duration of hospital stay has a non-zero tau-squared (2.444) alongside an I-squared of 0.0%: with large within-trial variances, appreciable between-trial variance still accounts for a negligible share of total variation. This is the expected behaviour of the two statistics, not a contradiction, and it is the reason tau-squared and the prediction interval are reported alongside I-squared throughout.

Baujat plot: contribution to heterogeneity against influence on the pooled estimate, weight gain velocity
Figure 15. Baujat plot: contribution to heterogeneity against influence on the pooled estimate, weight gain velocity.

3.8 Subgroup analyses

All three prespecified moderators were tested for every outcome with at least two trials per level (30 tests in total). No test was significant at the 5% level; the smallest interaction p-value was 0.114 (NEC (Bell stage >=2), definition of early fortification). Full results are in subgroup_analyses.csv.

Three cautions apply, and together they mean nothing in this section should be read as evidence of effect modification:

  1. Fortifier source and gestational-age stratum are collinear in this evidence base. The trials using a human milk-derived fortifier are exactly the trials restricted to extremely preterm infants, so the two moderators produce identical between-group statistics and cannot be distinguished.
  2. Subgroups contain one to five trials. Between-group tests use conventional random-effects standard errors because the Hartung-Knapp adjustment is unstable at two to three trials per level; this makes the tests anticonservative, and they still found nothing.
  3. The comparisons are between trials, not within them. No trial randomised infants to different definitions of "early", so every subgroup contrast is observational across trials and confounded by everything else that differs between them.

The subgroup analyses are reported because they were prespecified, and their uniform null result is itself informative only in the weak sense that the data provide no signal to pursue.

3.9 Sensitivity analyses

Four prespecified sensitivity analyses were run per outcome where applicable, plus leave-one-out influence analysis. Full results are in sensitivity_analyses.csv and influence_leave1out.csv; the primary outcomes are tabulated here.

Analysis Time to regain birth weight (d) Weight gain velocity (g/kg/d)
Primary (all trials, HK) -1.10 (-2.09 to -0.12, 4 trials) +0.30 (-0.99 to +1.59, 6 trials)
Excluding imputed SD -1.27 (-2.93 to +0.39, 3 trials) +0.34 (-7.06 to +7.75, 2 trials)
Excluding early-unfortified-comparator framing -1.12 (-2.89 to +0.65, 3 trials) +0.59 (-1.31 to +2.49, 4 trials)
Excluding trials at high overall risk of bias -0.79 (-2.82 to +1.24, 2 trials) -0.69 (-2.73 to +1.35, 3 trials)
Full-text-derived data only -0.97 (-1.59 to -0.36, 3 trials) not estimable

The birth-weight-regain result is fragile. It is the review's only interval excluding no effect, and it loses significance under two of the three applicable sensitivity analyses: excluding the trial with an imputed SD (-1.27 (-2.93 to +0.39, 3 trials)) and excluding trials at high overall risk of bias (-0.79 (-2.82 to +1.24, 2 trials)). The direction of effect is stable across every analysis and every leave-one-out subset — the point estimate stays between -1.27 and -0.97 days — but its statistical significance depends on retaining trials that the risk-of-bias assessment flags. This is why the outcome is rated low rather than moderate certainty.

Weight gain velocity is null under every analysis, and its interval widens dramatically when the four trials with an imputed SD are removed (+0.34 (-7.06 to +7.75, 2 trials)), which reflects how much of the apparent precision in that pooled estimate rests on reconstructed dispersion.

Duration of hospital stay changes sign between the primary analysis (+1.64 (-2.19 to +5.46, 6 trials)) and the imputed-SD exclusion (-2.29 (-5.02 to +0.43, 2 trials)), neither interval excluding no effect. No conclusion about hospital stay is supportable.

For the rare-event safety outcomes, restricting to trials not at high risk of bias leaves one or two trials and produces intervals so wide as to be uninformative — mortality 1.35 (0.0040 to 508.00, 2 trials) and NEC 0.87 (0.05 to 14.37, 3 trials). These are reported for completeness rather than interpreted.

3.10 Reporting bias

No formal small-study-effect test was performed for any outcome. The protocol set a threshold of 10 contributing trials before an Egger or Peters test is interpretable; the largest evidence base here is 7 trials reporting NEC. Funnel plots for the two primary outcomes and for NEC are presented as descriptive displays only — at four to six points, apparent asymmetry cannot be distinguished from sampling variation, and no visual judgement was made from them.

Funnel plot, time to regain birth weight (descriptive only)
Figure 16. Funnel plot, time to regain birth weight (descriptive only).
Funnel plot, weight gain velocity (descriptive only)
Figure 17. Funnel plot, weight gain velocity (descriptive only).
Funnel plot, NEC (descriptive only)
Figure 18. Funnel plot, NEC (descriptive only).

The structured assessment of reporting bias (reporting_bias_assessment.csv) therefore rests on what can be checked rather than on a test statistic:

Domain Finding Implication
Formal small-study-effect test Not performed for any outcome. The largest evidence base is NEC (Bell stage >=2) with 7 trials reporting the outcome and 5 contributing analysable data; the protocol (section 9) requires >=10 contributing trials before an Egger or Peters test is interpretable. No test statistic is reported anywhere in this review. Reporting bias cannot be assessed statistically for any outcome. The GRADE publication-bias domain is rated 'undetected but unassessable' throughout and no outcome is downgraded on it.
Funnel asymmetry (visual) Descriptive contour-free funnel plots are presented for the two primary outcomes (4 and 6 trials) and for NEC (5 analysable trials). With k of this size, apparent asymmetry cannot be distinguished from sampling variation and no visual judgement was made. Presented for completeness only; no evidence either way.
Outcome availability Across the 7 trials in quantitative synthesis, the 11 prespecified outcomes are reported by 3 to 7 trials each (median 5). Both primary outcomes are reported as a pair by 3 trials (Gupta 2025, Salas 2025, Shah 2016). Availability tracks outcome type rather than result direction: growth velocities and hospital-stay outcomes are widely reported, while time to regain birth weight (4 trials) and all-cause mortality (3 trials) are sparser. No trial was found to report an outcome narratively as non-significant while withholding the numbers. Availability is uneven but shows no direction-dependent pattern; residual outcome-level reporting bias cannot be excluded.
Prospective registration 5 of the 7 trials in synthesis report a registration identifier (Shah 2016 NCT01988792; Salas 2023 NCT04325308; Salas 2025 NCT05525585; Sullivan 2010 NCT00506584; Gupta 2025 CTRI/2020/07/026744), each verified against the source record. Siddiqui 2025 (preprint) and Wynter 2024 report none. Alizadeh Taheri 2017, narrative-only, also reports none. Prespecified-versus-reported outcomes could be compared for 5 of 7 trials; selective reporting cannot be excluded for the remaining 2.
Grey literature / unpublished The search covered Europe PMC preprint records and ClinicalTrials.gov. One unpublished preprint RCT (Siddiqui 2025, Research Square) was identified and included, with its non-peer-reviewed status carried into its risk-of-bias judgement. One eligible trial is available only as a paywalled publisher abstract (Alizadeh Taheri 2017); it reports no arm-level numbers, SDs, or event counts and so contributes to narrative synthesis only. No conference abstracts were retrievable in full. Grey literature was searched and one preprint recovered; a known eligible trial (80 infants) is nonetheless absent from every pooled estimate for want of extractable data.
Ongoing or unreported trials One registry record on fortification timing (NCT05251441) has no posted results and no linked publication, and is listed as awaiting classification. NCT04284280 (routine versus selective fortification) was excluded as the wrong comparison. No trial identified as complete was found to have gone unreported. One eligible-looking trial of unknown size and status remains unaccounted for.
Selective analysis reporting Salas 2023 reports both unadjusted and covariate-adjusted growth analyses; unadjusted values were used for pooling and the adjusted values recorded alongside in extracted_data.csv. No trial was found to have switched a declared primary outcome, though this could be checked only for the 5 registered trials. A registry-linked search on each registration identifier returned one companion publication, Tucker 2025 (secondary analysis of Salas 2023 reporting BPD severity), excluded as a non-independent report of an included trial. No included trial has an unaccounted companion report. Analysis-choice bias is unlikely to drive the pooled results but cannot be excluded for the 2 unregistered trials.
Within-trial data completeness 16 of the extracted continuous data points required an SD to be reconstructed from a reported SEM, IQR, or 95% interval (flagged imputed_sd in extracted_data.csv, affecting 4 trials). Two trials recorded zero events in both arms for NEC and two for late-onset sepsis; these carry no information for a risk ratio and are dropped from those pooled estimates, which is why the analysed and reporting trial counts differ for the two rare-event safety outcomes. Imputation and double-zero exclusion are recorded per data point; sensitivity analyses excluding imputed-SD trials are reported in sensitivity_analyses.csv.

The material points are that 5 of 7 trials in synthesis are registered, so prespecified-versus-reported outcomes could be compared for most but not all of them; that one eligible trial (80 infants) is absent from every pooled estimate because its full text could not be obtained; and that one registry record has no results and no linked publication. Outcome availability across trials tracks outcome type rather than result direction, which is reassuring but not conclusive.


4. Certainty of evidence (GRADE)

GRADE summary of findings
Figure 19. GRADE summary of findings.

Certainty was rated per outcome against the protocol's minimally important differences. The distribution is 7 very low, 3 low, 1 moderate across 11 outcomes. Every outcome was downgraded at least once for risk of bias, since no trial was at low risk overall. Written reasons for every domain judgement are in grade_summary_of_findings.csv.

Three features of the rating deserve stating explicitly:

  1. No outcome was downgraded for publication bias, but neither was any outcome cleared of it. With no outcome reaching the testing threshold, the domain is recorded as "undetected but unassessable" throughout, which is a statement about the evidence base rather than about the absence of bias.
  2. The only moderate-certainty outcome is time to full enteral feeds, downgraded once for risk of bias and not at all for imprecision, inconsistency or indirectness. It is the strongest result in the review, and it is a null.
  3. The one outcome with a significant effect is rated low certainty. Time to regain birth weight is downgraded for risk of bias and for imprecision — the interval's lower bound only just reaches the 2-day minimally important difference while its upper bound is close to no effect — so the finding should be treated as a signal for future trials, not as a result to act on.

5. Discussion

5.1 Summary of findings

Across 7 randomised trials and 706 infants, early fortification of human milk was associated with regaining birth weight approximately one day sooner (MD -1.10 (95% CI -2.09 to -0.12) days, low certainty), and with no detectable difference in any other prespecified outcome: weight gain, linear and head growth, growth failure at discharge, time to full feeds, feed intolerance, NEC, mortality, sepsis, or hospital stay. The birth-weight-regain finding does not survive exclusion of high-risk-of-bias trials or of the trial with an imputed SD, and its magnitude is below the prespecified minimally important difference of 2 days.

The safety picture is the practically important one. The concern that motivated deferring fortification is NEC, and the pooled estimate (RR 0.96 (95% CI 0.32 to 2.90)) is centred almost exactly on no effect with an interval spanning a two-thirds reduction to a near-tripling. 19 events in 516 infants cannot resolve a 25% relative difference. The correct reading is not "early fortification is safe" but "these trials are too small to detect the harm they were concerned about, and they show no sign of it."

5.2 Comparison with previous reviews

CD013392 (search to August 2019) pooled two trials and 237 infants and concluded that the evidence was insufficient. This review adds 5 trials and roughly triples the number of infants, and reaches a materially similar conclusion for every outcome except time to regain birth weight, where the accumulated data now produce an interval that excludes no effect under the primary analysis but not under sensitivity analysis. The direction of travel is that a threefold increase in evidence has not changed the qualitative answer, which suggests the remaining uncertainty is not going to be resolved by another trial of 60 to 150 infants.

5.3 Limitations

Of the evidence. No trial was at low risk of bias overall. Masking is impossible for the intervention and largely absent, which matters most for the clinician-assessed outcomes (feed intolerance, the decision to withhold feeds, time to full feeds). Trials are small (26 to 150 infants per trial) and the rare-event outcomes have too few events to be informative. Definitions of "early" and "delayed" vary substantially — from the first feed to day 7, and from 70 mL/kg/day to day 14 — so the pooled comparison averages over materially different interventions. Two trials document protocol drift or a narrowed contrast that biases their contribution toward the null.

Of this review. Screening and extraction were performed by a single assessor without duplication, so no inter-rater agreement can be reported and selection errors cannot be excluded. Trial authors were not contacted, which is why one eligible trial contributes nothing to the pooled estimates and why 16 data points across 4 trials required an SD to be reconstructed rather than obtained. Two eligible reports were available only as abstracts: one (Siddiqui 2025) was extracted from a preprint abstract and does contribute to the pooled estimates, and the other (Alizadeh Taheri 2017) reports no extractable numbers at all. The review is not prospectively registered. These are declared in the protocol rather than discovered afterwards, but they are real constraints on how much weight the results can bear.

5.4 Implications for practice

The evidence does not support a general recommendation to change practice in either direction. It does undercut the specific rationale for routinely deferring fortifier to 100 mL/kg/day: across 7 trials there is no signal of increased NEC, feed intolerance or sepsis with early fortification, and the one outcome measured precisely enough to be informative (time to full enteral feeds, moderate certainty) shows no delay. Units that already fortify early have no evidence of harm to act on; units that defer have no evidence of benefit from deferring. Given very low to low certainty for every safety outcome, this is a case for local protocol consistency and prospective audit rather than for a strong recommendation.

5.5 Implications for research

A trial adequately powered for NEC is the missing piece and is not feasible as a single-centre study: at the control-arm risk observed here (3.9%, 10 events in 256 control infants), detecting a 25% relative reduction with 80% power at a two-sided alpha of 0.05 requires about 5,400 infants per arm — roughly 10,800 in total, or 15 times the entire evidence base assembled here. The realistic paths forward are a prospective individual-participant-data collaboration across the existing and ongoing trials, which would also allow the effect-modification questions in section 3.8 to be asked within trials rather than across them, and a registry-based or cluster-randomised design large enough for the rare outcomes. Future trials should also standardise on a single definition of early initiation, report arm-level SDs for every growth outcome, and register and report a fixed primary outcome — three of the review's largest sources of uncertainty are avoidable reporting problems rather than irreducible biology.


6. Data availability and reproducibility

Every number in this report is generated from the files below by report.py, which reads the analysis outputs directly; no value is typed into the prose.

File Contents
protocol.md Protocol, fixed before extraction
screening_log.csv Every screening decision with its reason
extracted_data.csv Long-format extracted data, one row per trial per outcome, with source and access provenance per datum
study_characteristics.csv Trial-level characteristics
risk_of_bias.csv RoB 2 judgements with written justification per domain per trial
analysis.R Meta-analysis, heterogeneity, subgroups, sensitivity analyses, forest/funnel/Baujat plots
figures.py PRISMA diagram, characteristics table, traffic light, GRADE figure
report.py This report
build_html.py Renders this report to a single self-contained report.html with all figures embedded
pooled_results.csv Primary pooled estimates (Hartung-Knapp)
pooled_results_classicCI.csv Conventional Wald intervals, with available and analysed trial counts
heterogeneity.csv tau-squared, I-squared with CI, Q, prediction intervals
subgroup_analyses.csv All 30 prespecified subgroup tests
sensitivity_analyses.csv Prespecified sensitivity and leave-one-out analyses
influence_leave1out.csv Leave-one-out estimates with recomputed heterogeneity
grade_summary_of_findings.csv GRADE domain judgements with written reasons
reporting_bias_assessment.csv Structured reporting-bias assessment
narrative_only_studies.csv Eligible trials excluded from pooling, with reason

To reproduce: Rscript analysis.R, then python figures.py, then python report.py, then python build_html.py. The last step produces report.html, a single file with every figure embedded and no external dependencies; use the browser's print function for a paper copy.


7. Included trials