78% mortality in failed weaning, the Hypotension Prediction Index explained away, the surgical establishment on TAVI durability, and bigger annuloplasty rings failing later
Four items, all published 9 September. Two of them are reinterpretations rather than new data — and both change what an existing literature means.
Four items, every one published 9 September, all with deposited abstracts read in full. Two are reinterpretations of existing literature rather than new data, and both are more consequential for that — one explains away a monitoring technology’s apparent benefit, the other reopens a question the field had treated as closed.
1. WEAN SAFE — failed weaning carries 78% ICU mortality, and the first separation attempt predicts the decision to stop
Journal · Published: Intensive Care Medicine, 9 September 2026 — published yesterday Caldecott, Zhu, McNicholas, Rezoagli, Taran, Pham, Heunks, Bellani, Brochard, Simpkin, Laffey, and the WEAN SAFE Investigators · NCT03255109
Secondary analysis of the WorldwidE AssessmeNt of Separation of pAtients From ventilatory assistancE cohort, looking specifically at patients who enter weaning and fail.
Of 5,664 patients:
- 1,270 (22.4%) never underwent a separation attempt at all
- 4,394 (77.6%) had at least one, of whom 3,730 (65.9%) weaned successfully and 664 (15.1%) had failed weaning at day 90
- Failed-Wean patients had the longest duration of invasive ventilation and longest ICU stay of any group
- Compared with Successful-Wean, they underwent more separation attempts but fewer attempted extubations, and more reintubations and tracheostomies
- ICU mortality 78% in Failed-Wean versus 2% in Successful-Wean
Three phenotypes, distinguished by separation-attempt count and the timing of withdrawal or withholding of life-sustaining therapy (WLST):
- Phenotype A — single separation attempt plus WLST
- Phenotype B — single separation attempt, no WLST
- Phenotype C — more than one separation attempt
And the finding the authors flag for further investigation: a failed first separation attempt was strongly associated with WLST decisions, and a WLST decision was associated with a higher hazard of ICU mortality and a shorter median time to ICU death.
Interpretation. The 78% versus 2% mortality gap is the number to carry, but it is not the interesting part — failing to wean is obviously a marker of a dying patient. The interesting part is the structure the phenotypes expose, and it should be read carefully because it is easy to over-read.
A failed first separation attempt predicted the decision to withdraw. That is a finding about clinician behaviour, not about lung or diaphragm physiology. Phenotype A — one attempt, then WLST — is a group in which the weaning trajectory effectively ended after a single trial. The authors are appropriately careful: they say the relationship “deserves further investigation”, not that it is causal. But the honest reading is that the first SBT is functioning as a prognostic test that informs end-of-life decisions, which is a role it was never validated for.
There are two mutually exclusive explanations and this design cannot separate them. Either the first failed attempt is genuinely identifying patients who will not recover — in which case clinicians are reading it correctly — or it is triggering a self-fulfilling trajectory, where an early failure shifts the team’s expectations, attempted extubations become less frequent (which the data show: fewer extubation attempts despite more separation attempts), and the patient is managed toward death. That ambiguity is exactly the shape of the WLST problem in neuroprognostication after cardiac arrest, and it is no more tractable here.
The 22.4% who never had a separation attempt at all is the other number worth pausing on. More than one in five ventilated patients in a worldwide cohort never reached the point of a weaning trial. Whatever is happening in that group, it is not weaning failure — it is a different and largely unstudied population.
Practically: this is a reason to be explicit in handover about whether a failed SBT is being treated as physiology or as prognosis. Those are different claims and the first attempt conflates them.
Read the paper · PMID 42714473
2. The Hypotension Prediction Index — the benefit tracks how hard you treat, not what you predicted
Journal · Published: Anesthesiology, 9 September 2026 — published yesterday Vistisen, Novotny, Mukkamala, Enevoldsen, Jacquet-Lagrèze, Elbers
The setup, in the authors’ own framing. HPI is marketed for proactive prevention of hypotension, and meta-analyses report less hypotension with HPI-guided management versus a MAP-alert-at-65 strategy. But validation studies suggest HPI does not outperform simple MAP-based prediction, and HPI alerts fire when MAP is around 70–75 mmHg. So existing HPI trials may be comparing two different MAP treatment thresholds — 70–75 against 65 — rather than prediction against no prediction. The hypothesis: reductions in hypotension are associated with greater treatment intensity.
Twenty randomised trials, 2,342 patients; 18 reported both hypotension and treatment outcomes.
- Among the 13 trials that reduced hypotension, 9 (69%) reported significantly greater haemodynamic treatment — fluids, vasopressors or inotropes — in the HPI arm
- Among the 5 trials that failed to reduce hypotension, none reported increased treatment intensity
- Association between hypotension reduction and treatment intensity: P = 0.029
- For trials reporting time-weighted average MAP < 65 mmHg, the TWA difference correlated with treatment intensity category: P < 0.001
Interpretation. This is the most intellectually satisfying paper of the day, and the one most likely to change a purchasing decision.
The argument is simple and close to airtight in structure: if the intervention arm got more fluid and more vasopressor, and the intervention arm had less hypotension, you have not demonstrated that prediction works. You have demonstrated that treating earlier works — which nobody disputed. Every trial that reduced hypotension without increasing treatment intensity would be evidence for the algorithm; there were four of those, against nine where treatment intensity rose. And the complete absence of increased treatment in the five null trials closes the loop neatly: where treatment did not intensify, nothing happened.
Two things this paper does not establish, and it is careful about both. It does not show HPI is useless — an alert that reliably prompts earlier treatment has value even if its predictive content is no better than a MAP threshold, because prompting is itself an intervention. And it is an association across trials, not a within-trial mediation analysis; twenty trials and a 2×2 table cannot prove mechanism.
But the practical conclusion stands: the appropriate comparator for HPI was never MAP < 65. It was MAP < 75. If you want the benefit these trials show, you may be able to get most of it by treating hypotension earlier with the monitor you already own. That is a testable claim, and it is notable that a decade of HPI trials did not test it.
Sits naturally beside the archive’s running theme from the renal perfusion phenotype cohort — perfusion not pressure — with a sharper version of the same caution: be clear what your target variable is actually doing, and what the comparator was.
Read the paper · PMID 42713963
3. TAVI in low-risk patients — the surgical establishment calls for caution on durability
Journal · Published: The Annals of Thoracic Surgery, 9 September 2026 — published yesterday Sousa Uva, Milojevic, Marin-Cuartas, Kaul, De Caterina, Redberg, Heuts, Doenst, Dayan, Siepe, Badhwar, Falk, Borger, Myers, Sadaba
A review, not new trial data — but the author list is most of the European and American cardiac surgical leadership, and it is the document that matters this week in cardiothoracic surgery.
The concern: TAVI use is expanding in low-risk patients “including populations underrepresented in pivotal trials and often beyond established guideline recommendations.” What the extended follow-up shows:
- Evolut Low Risk, 6–7 years — no statistically significant difference in the primary composite of death or disabling stroke, but signals of later mortality accrual, more myocardial infarction, and higher aortic valve reintervention with TAVI
- PARTNER 3, 7 years — outcomes “appeared broadly similar” on reported composites, but interpretation limited by heterogeneous endpoint construction, non-prespecified hierarchical analysis, incomplete follow-up, and the influence of post hoc vital-status ascertainment on late mortality estimates
- PARTNER 2A (intermediate risk), 10 years — lower survival and higher aortic valve reintervention with TAVI than SAVR
- With UK-TAVI, meta-analyses and large observational data, these “underscore uncertainty about durability and reintervention burden”
The call: a more rigorous and transparent long-term evidence framework before further expanding TAVI, particularly in younger, low-risk patients, with decisions individualised in a structured Heart Team framework including explicit discussion of expected survival, anatomical suitability, valve durability, reintervention options and lifetime management.
Interpretation. Read this with the bias declared on both sides. This is cardiac surgeons writing in a cardiac surgical journal about a procedure that has taken work away from cardiac surgeons — and the substance is still largely right.
The PARTNER 3 critique is the strongest part and the least arguable. “Post hoc vital-status ascertainment” influencing late mortality estimates is a serious methodological objection, not a turf complaint: if you go looking for deaths after the fact and find them unevenly between arms, your late mortality comparison is compromised. So is non-prespecified hierarchical analysis. These are the kinds of problems that make a “broadly similar” conclusion unreliable in either direction.
PARTNER 2A at ten years is the number that should change behaviour, and it is about intermediate-risk patients — lower survival and more reintervention with TAVI. Ten-year data is what a 65-year-old needs and what the low-risk trials cannot yet provide. The Evolut signal at 6–7 years — mortality accruing later, more reintervention — is the direction you would expect if durability diverges over time, and the honest statement is that the primary endpoint was null and the secondary signals are suggestive.
This is the fourth time this archive has landed on the same structural problem: a guideline or practice shift resting on trials whose follow-up is shorter than the patient’s remaining life. The IACTS position statement on the CABG downgrade, the Medicare survival analysis urging reevaluation of guidelines, the death of David Taggart, and now TAVI durability. The disputes are not the same dispute, but they share a shape — and for the perioperative team, the practical consequence is identical: the patients who arrive for surgery are increasingly the ones for whom the less invasive option already failed.
One caution on the paper itself: it is a review with a stated position, and a reader should go to the Evolut and PARTNER publications for the numbers rather than take the characterisations here as neutral.
Read the review · PMID 42716272
4. Annuloplasty ring size — bigger rings fail later, with a threshold above 34 mm
Journal · Published: The Annals of Thoracic Surgery, 9 September 2026 — published yesterday Malik, Rheault-Henry, Cubas Llalle, Chu (single-surgeon series)
The premise is that larger rings are generally preferred in degenerative mitral disease — better haemodynamics, less systolic anterior motion, restored annular geometry. This series tests that against long-term durability.
503 patients, mitral valve repair for degenerative mitral regurgitation, 2008–2024, one surgeon. Primary endpoint repair failure — reintervention, recurrent MR, or mitral stenosis. Stratified by ring size quartile; Cox models with a time-dependent spline interaction to capture late effects. Median follow-up 5.3 years (IQR 2.8–8.2).
- Failure by quartile: ≤32 mm 2.0%, 33–34 mm 3.2%, 35–36 mm 5.7%, >36 mm 7.3%
- Ring size an independent predictor: HR 1.06 per mm (1.01–1.10), P < 0.01
- Time-dependent analysis: increased hazard beyond 11.7 years
- Optimal cut-point for long-term failure: >34 mm
Interpretation. The design choice is what makes this worth reading: a time-dependent spline rather than a single proportional hazard. The effect did not appear until beyond 11.7 years, which is longer than the median follow-up of most mitral repair series — and longer than this one. That is simultaneously the paper’s contribution and its main limitation.
The contribution: a late-emerging hazard is invisible to conventional analysis, and this is a plausible mechanism for why repair durability curves diverge in the second decade. A 1.06 hazard ratio per millimetre is small, but over the 6 mm spanning the quartiles it compounds, and the crude failure rates — 2.0% to 7.3% — are a more than threefold difference.
The limitation: a hazard that begins at 11.7 years, estimated from a cohort with median follow-up of 5.3 years, rests on the minority of patients followed that long. Those are the earliest-operated patients, from 2008 onward, which confounds ring size with era, technique and surgeon experience. A single-surgeon series controls for the variability between surgeons at the cost of being one surgeon’s practice.
What it does not tell you is the trade-off. Larger rings were preferred for reasons — systolic anterior motion and mitral stenosis being the feared alternatives — and this paper reports repair failure including mitral stenosis as a composite, so it cannot tell you whether smaller rings simply move the failure mode. “More physiologic sizing” is the authors’ phrasing, and it is carefully non-specific.
For the cardiac anaesthetist running the post-repair TOE, the practical note is narrow but real: a ring above 34 mm in degenerative disease now has a published association with late failure, which makes the completion assessment of residual MR worth recording carefully rather than waving through. That is the same lesson as yesterday’s aortic root preservation series, where intraoperative residual mild AR carried a hazard ratio of 4.08 for late recurrence. Two papers in two days saying the end-of-case echo predicts the next decade.
Read the paper · PMID 42716273
Notes for the next run
- Seen and not pursued, worth a slot on a quiet day: Shock, 9 September — “Early antithrombin depletion after trauma is associated with biomarkers of endotheliopathy and systemic protein loss” (Matsumoto et al., DOI 10.1097/shk.0000000000002935, PMID 42713784). Full abstract available and read: 100 trauma patients, reduced AT activity in 32%, adjusted OR 9.42 for shock and 12.20 for DIC; AT tracked albumin and syndecan-1 rather than thrombin-antithrombin complex, arguing the depletion is endothelial leak and protein loss rather than thrombin consumption. Single centre, n=100, wide confidence intervals (the shock OR reaches 109) — interesting mechanism, weak estimates.
- No abstract deposited: Anesthesia & Analgesia, 9 September — “Anti-Xa Activity and Global Hemostatic Response After Prophylactic Enoxaparin in Term Pregnancy: A Prospective Matched-Cohort Study” (Nguyenová et al., DOI 10.1213/ane.0000000000008296, PMID 42715360). Directly relevant to the neuraxial timing intervals in the ESAIC/ESRA antithrombotic guideline — retry for the abstract in a few days, or it needs a PDF.
- Also seen, not pursued: an A&A J-PEDIA registry analysis of extreme weight-for-age and airway adverse events at induction (abstract not deposited), the MATTERHORN final analysis of perioperative durvalumab in gastro-oesophageal adenocarcinoma (Lancet, oncology rather than perioperative medicine), and a large volume of EJA and A&A correspondence.