Drug allergy labels and worse surgical outcomes in 13,646 patients, low ETCO₂ as an independent mortality marker, code-status limitations and early death, and an STS risk model that failed its own validation
Four items. Three are large observational cohorts, and all three are unusually careful about the difference between what they measured and what it means.
Four items — one published today, three yesterday, all with deposited abstracts read in full. Three are large observational cohorts, and what links them is that each is explicitly careful about the gap between the association it found and the inference someone will draw from it. Two of them say so in their own conclusions.
1. SAPPHIRE — a drug allergy label is associated with worse surgical outcomes in 13,646 patients
Journal · Published: British Journal of Anaesthesia, 10 September 2026 — published yesterday Savic, Dias, Vairale, Begum, Khan, Fowler, Kaura, Watson, Littlejohns, Pearse, Abbott and the SAPPHIRE study investigators · ISRCTN15775657
Prospective multicentre observational study across 21 UK NHS hospitals. Adults undergoing primary hip or knee replacement, internal fixation of a closed long bone fracture, colorectal resection, TURP or TURBT, Caesarean delivery, or hysterectomy. Recent antibiotic use excluded. Primary outcome a 30-day composite: all postoperative infections, anastomotic leak, ARDS, myocardial infarction, bleeding, pulmonary embolism, stroke, antimicrobial side-effects, death.
The premise, stated at the top of the abstract: one in four surgical patients carries a drug allergy label, of which an estimated 90% are incorrect.
13,646 patients; 3,924 (29%) carried at least one label.
- Primary composite: 25% (989/3,924) labelled vs 20% (1,926/9,722) unlabelled — OR 1.21 (1.10–1.34), P < 0.001
- Surgical site infection: 9% vs 8% — OR 1.19 (1.03–1.38), P = 0.018
- Any postoperative infection: 19% vs 15% — OR 1.24 (1.11–1.38), P < 0.001
- Allergic drug reactions: 31/3,924 vs 29/9,722 — OR 3.00 (1.77–5.09), P < 0.001
- No increase in mortality
Interpretation. The mechanism the authors propose — “avoidance of first-choice drug therapies can lead to worse postoperative outcomes” — is almost certainly the main story, and the infection signal is where it shows. A patient labelled penicillin-allergic gets a second-line prophylactic antibiotic, and second-line prophylaxis is less effective. An OR of 1.19 for surgical site infection in a cohort this size is consistent with exactly that.
But the 3.00 odds ratio for allergic drug reactions is the finding that should make you stop, because it runs against the comfortable version of this story. The comfortable version is “these labels are 90% wrong, so they are noise and the harm is entirely iatrogenic avoidance.” If that were the whole truth, labelled patients would not have three times the rate of actual allergic reactions. Some of these labels are real, or at least mark patients genuinely prone to drug reactions — and a de-labelling programme that treats every label as spurious will hurt somebody.
The honest reading of both findings together: labels cause harm through avoidance, and labels also carry real signal, and the same cohort shows both. That is an argument for investigating labels rather than for either ignoring or obeying them.
Two limits worth stating. This is observational, and a drug allergy label is not randomly distributed — labelled patients differ in healthcare contact, comorbidity and documentation intensity, all of which independently predict complications. And the absolute differences are small: 25% versus 20% on a broad composite, 9% versus 8% for SSI. No mortality difference. This is a real effect of modest size, not a crisis.
Practically, the actionable item is narrow and worth doing: the patient labelled penicillin-allergic having a colorectal resection is the one to interrogate preoperatively, because that is where second-line prophylaxis and a high baseline infection rate meet. It connects to the antibiotic thread this archive has been building — the prolonged-infusion β-lactam focused update and the IDSA sepsis position paper both optimise β-lactam delivery, which does nothing for the 29% of patients not getting a β-lactam at all.
Read the paper · PMID 42722596
2. Low intraoperative ETCO₂ and postoperative mortality — 185,455 patients, independent of hypotension and minute ventilation
Journal · Published: Anesthesiology, 10 September 2026 — published yesterday Huz, Lamer, Bourgeois, Cirenei, Moussa, Chazard, Tavernier (Lille)
Retrospective cohort, adults undergoing noncardiac surgery under general anaesthesia with mechanical ventilation, 2010–2020, single tertiary centre. The question is specifically about confounding: low ETCO₂ has been linked to mortality before, but is that independent of intraoperative hypotension (a strong predictor in its own right) and of minute ventilation (the primary determinant of ETCO₂)?
185,455 patients; in-hospital mortality 0.85%.
- Lower mean intraoperative ETCO₂ was nonlinearly associated with mortality: adjusted OR 1.63 (95% CI 1.36–1.86) per 5 mmHg decrease from the median
- Independent of minute ventilation and hypotension
- No significant interaction between ETCO₂ and hypotension (P = 0.19)
- Robust across sensitivity analyses
- Authors’ conclusion: ETCO₂ “provides prognostic information beyond arterial pressure alone and may be a valuable marker for postoperative risk stratification”
Interpretation. The design is the point. Adjusting for minute ventilation is what separates this from the earlier ETCO₂ literature, because without it you cannot tell whether low ETCO₂ means poor perfusion or just hyperventilation by the anaesthetist. Having adjusted for it, the association persists — which leaves low ETCO₂ as a plausible marker of low cardiac output or increased dead space rather than an artefact of ventilator settings.
Note precisely what the authors claim: a marker for risk stratification. Not a target. The abstract does not say to raise ETCO₂, and the distinction matters enormously. A retrospective association between a physiological variable and death does not imply that moving the variable moves the outcome — that is the error the entire goal-directed therapy literature has spent twenty years making.
And read it directly against yesterday’s Hypotension Prediction Index paper, which appeared in the same journal one day earlier. That analysis showed the apparent benefit of HPI-guided care tracked treatment intensity, not prediction — covered here yesterday. Two papers, consecutive days, one journal, making complementary points: a monitored variable can carry real prognostic information (ETCO₂) and still not be a useful target, and a variable you successfully move can improve outcomes for reasons unrelated to the monitor that prompted you. Together they are a compact education in why perioperative monitoring trials are hard.
The limits: single centre, retrospective, a decade of changing practice, and 0.85% mortality means the entire signal rests on roughly 1,600 deaths spread across a very large denominator. “Nonlinearly associated” is also doing real work — the OR per 5 mmHg is an average over a curve, and where on that curve the risk actually climbs is not in the abstract.
Read the paper · PMID 42720197
3. Code-status limitations continued into theatre — 2.9% versus 1.2% three-day mortality, and the authors say what it means
Journal · Published: Annals of Surgery, 11 September 2026 — published today Allen, Streid, Lilley, Cauley, Bernacki, Reich, John, Hepner, Bader, Shah (Brigham)
Retrospective cohort, March 2024 to June 2025, 5 hospitals in one academic health system. Adults presenting for a procedure under anaesthesia with a code status other than full code. Exposure: whether the limitation was reversed to full code perioperatively or not. Primary outcome all-cause mortality within 3 days.
2,833 patients, none excluded. Median age 79 (IQR 70–86); 59% women.
- 2,323 (82%) reversed to full code perioperatively; 510 (18%) remained not full code
- 44 patients (1.6%) died within 3 days
- 3-day mortality: 2.9% (15/510) remaining not full code vs 1.2% (29/2,323) reversed
- Adjusted for age, sex, ASA status, operative stress score and race: adjusted OR 2.18 (95% CI 1.14–4.16); adjusted absolute risk difference 1.43% (0.06–2.95%)
- Of those who died within 3 days and remained hospitalised, 34 of 42 (81%) died after transition to comfort-focused care
- Invasive haemodynamic monitoring was more common in those remaining not full code — adjusted absolute risk difference 3.74% (0.91–6.93%)
And the authors’ own closing sentence, which is the reason to read this paper:
This pattern reflects downstream clinical trajectories and treatment decisions rather than missed opportunities for perioperative rescue.
Interpretation. This is the most ethically careful paper of the day, and the care is in the conclusion rather than the numbers.
A naive reading of “OR 2.18 for death if you keep the DNR” is that patients died because resuscitation was withheld — that the limitation killed them. The authors explicitly reject that, and the 81% figure is why they can: of those who died, four in five died after a transition to comfort-focused care, not during an un-attempted resuscitation in theatre. These were patients on a dying trajectory whose code status was an accurate reflection of that trajectory and of their own stated preferences. The code status was a marker of prognosis, not a cause of death.
The invasive monitoring finding is the genuinely surprising one and cuts against the stereotype. Patients who kept their limitation received more invasive haemodynamic monitoring, not less. That is the opposite of the feared “DNR means do not treat” drift — it suggests teams were escalating monitoring precisely because they could not escalate resuscitation, which is a coherent and arguably correct response.
The 82% reversal rate is the number to sit with. Four in five patients with a pre-existing code-status limitation had it reversed to full code for theatre. Some of that is appropriate and consented — an anaesthetic makes certain reversible arrests likely and easy to treat. But 82% is high enough to raise the question of whether these were individual conversations or a default, and that is a question about consent rather than outcome. Both ASA and the Royal College have guidance that perioperative DNR suspension should be a discussion, not an automatic policy.
It also pairs with a paper this archive noted but did not pursue: Regional Anesthesia and Pain Medicine on 4 September published “Do-not-resuscitate status and regional analgesia utilization for fracture pain: a comfort paradox?” — patients with DNR orders receiving less regional analgesia. Same underlying worry, opposite direction, and a reminder that code status leaks into decisions it should not touch.
Limits: single health system, 15 deaths in the exposed group, and the confidence interval on the absolute risk difference reaches 0.06% — the lower bound is essentially nothing.
Read the paper · PMID 42723113
4. Adult congenital cardiac reoperations — an institutional risk model fails in the STS database, and the rebuild is the finding
Journal · Published: The Annals of Thoracic Surgery, 10 September 2026 — published yesterday Griffeth, O’Sullivan, Dearani, Stephens, Todd, Egbe, Connolly, Burchill (Mayo)
External validation of an institutional risk model for reoperative adult congenital heart disease surgery, tested in the Society of Thoracic Surgeons Adult Cardiac Surgery Database, July 2017 to December 2023. Outcome: composite of operative mortality plus major morbidity (mechanical circulatory support, dialysis, unplanned noncardiac reoperation, neurologic deficit, cardiac arrest). Logistic regression and machine learning, with Shapley additive explanations for predictor importance.
- The outcome was nearly twice as common nationally: 16.7% in STS-ACSD vs 8.8% in the institutional cohort
- Attributed in part to urgent/emergent procedures — 26.6% nationally vs 5.4% institutionally
- The institutional model’s discrimination was attenuated, with systematic underestimation of risk
- A de-novo model from the 15 most influential STS variables achieved cross-validated AUROC 0.74 (extreme gradient boosting) and 0.73 (penalised logistic regression)
- Those 15: aortic procedure, status, haematocrit, ejection fraction, creatinine, age, white cell count, hypertension, connective tissue disorders, platelets, symptoms, reoperation number, heart failure, primary payor, BMI
Interpretation. The headline is a negative result reported properly, and that is rarer than it should be.
A model built at a high-volume centre systematically underestimated risk nationally, and the reason is in the numbers: 5.4% urgent/emergent at the referral centre versus 26.6% nationally. The institutional cohort was an elective, planned, selected population. Its model learned that world and then met a different one. This is the single most common way risk models fail, and it is worth naming because a clinician handed a calculator rarely knows which population it was trained on.
AUROC 0.73–0.74 for the rebuilt model is modest and should be said plainly. That is fair-to-moderate discrimination — better than nothing, not good enough to drive an individual operative decision on its own. Note also that extreme gradient boosting beat penalised logistic regression by 0.01, which is to say not at all. After all the machine learning, a regression model did the same job; the authors’ framing of “clinically interpretable risk models integrating ML and regression” is a graceful way of saying the ML did not add much.
“Primary payor” appearing among the 15 most influential predictors deserves a comment. In a US database, insurance status is a proxy for access, timing of referral, and whether a patient arrives elective or emergent — which makes it both genuinely predictive and a measure of something other than physiology. A risk model that performs better because it knows who is insured is telling you about a health system, not about a heart.
For the perioperative team the takeaway is the one the paper states directly: reoperative ACHD surgery remains high risk nationally — a 16.7% composite morbidity and mortality rate — and the population reaching theatre is increasingly adult, increasingly reoperative, and frequently urgent. It sits alongside the AHA statement on residual lesions after paediatric cardiac surgery sent on 8 September: the children in that document become the adults in this one.
Read the paper · PMID 42722265
Notes for the next run
- A null worth recording, not reporting: Anesthesia & Analgesia, 10 September — dexamethasone 8 mg versus ondansetron 4 mg as first-line antiemetic after Caesarean delivery (Berger et al., DOI 10.1213/ane.0000000000008321, PMID 42721472). Double-blind RCT, 100 enrolled and 95 completed, spinal with 150 µg morphine and an enhanced-recovery protocol. No difference in total medications for nausea, pain or pruritus at 24 h (0.23 vs 0.40 per patient, P = 0.337), and no difference in pain despite prior work favouring dexamethasone — which the authors attribute to low pain scores and little breakthrough pain in their cohort, i.e. a floor effect. Honest negative in a small trial; pairs with the 5th PONV consensus guidelines.
- Unretrieved from 9 September, retry: the A&A anti-Xa study after prophylactic enoxaparin in term pregnancy and the A&A J-PEDIA airway registry analysis both still have no abstract deposited in Europe PMC. Both remain on the outstanding list.
- Read and recorded, just short of a slot: Annals of Thoracic Surgery, 10 September — “Real-world adherence to guideline-directed imaging surveillance in chronic aortic dissection and thoracic aortic aneurysm: a scoping review” (Jacobson et al., DOI 10.1016/j.athoracsur.2026.08.021, PMID 42722267). Abstract deposited and read in full. PRISMA-ScR scoping review, six databases through May 2026, US and Canadian studies only. 13 retrospective studies: 1,344 Type A dissections, 865 Type B (357 treated invasively at index admission, 508 initially medically), 533 repaired and 29,185 unrepaired thoracic aortic aneurysms. Adherence to guideline-directed imaging surveillance was generally low and variable; surveillance was associated with detecting adverse radiographic findings and with higher intervention rates. On survival the results scatter in both directions — two studies associated surveillance with lower mortality, three found nonsignificant trends toward lower mortality, and two associated greater surveillance with higher mortality. Mostly single-centre chart reviews with varying definitions and intervals, and few datasets able to capture imaging done outside the primary health system — so the two studies showing higher mortality with more surveillance are most plausibly confounding by indication rather than harm. Worth having in mind against the EACTS/STS aortic organ guidelines, still partly unread on the outstanding list: those guidelines set the surveillance intervals this review finds are widely not followed, and the evidence that following them saves lives is — on this synthesis — not yet there.
- Also seen, out of scope: an AJRCCM RCT of high-intensity interval training in interstitial lung disease, two Annals of Surgery pieces on bariatric revision and emergency general surgery benchmarking, and an Annals of Thoracic Surgery analysis of Commission on Cancer Standard 5.8 adherence and nodal upstaging.