AI Foundry lab · Incident and lesson · 30 September 2026
A PCIe link retrain hung aifoundry1
Never retrain, re-speed or reset a card's PCIe link on a running host. At 14:41 PDT a test of aifoundry1 card 0's link, run as root, asked its root port to retrain at 8 GT/s, and the whole host stopped. It stayed down for 26 minutes, until Roman power-cycled the lab at about 15:07. All three machines came back on kernel 7.0.0-34, and all four cards work.
What happened
aifoundry1 card 0's PCIe link had logged about one corrected receiver error a second at 16 GT/s since the host booted on 18 September. To tell a marginal signal from a bad lane or slot before an on-site visit, our lab plan proposed a “Gen3 test” (U25): set the target speed of card 0's root port, 0000:00:01.0, to 8 GT/s and retrain the link for ten minutes, with card 0 idle and its lock held, then set it back. The plan said the worst case was card 0 dropping off the bus until the next reboot.
The owner ran it as root at 14:40:21. Two readings at 16 GT/s a minute apart went through: the root port's corrected receiver-error count rose from 1,146,936 to 1,147,012 (76 in 60 s), and the card's own count stayed at 0. The next command, at 14:41:21, wrote Link Control 2 (target speed) and set the Retrain Link bit with setpci. Nothing came back. Within twenty seconds aifoundry1 left Tailscale and stopped answering ARP on the LAN; a read-only command of ours had run there at 14:41:20. Every session on it, both cards, its CI runner and its /tmp went with it.
There is no console or out-of-band access to the lab machines, so nothing could restart it remotely. Roman power-cycled the lab in person at about 15:07; aifoundry1, aifoundry2 and aifoundry3 all rebooted within a minute of each other.
After the power cycle (checked read-only at 15:18–15:20)
- All three hosts run kernel 7.0.0-34 with
et_soc10.20.0, loaded at boot (on aifoundry1 for the first time since its fix of 25 September), boot to the text target, and reportrunningwith no queued jobs. - All four cards are on the bus with their nodes; every link is at 16 GT/s x8; nobody holds a card.
- aifoundry1 card 0's driver error counters are all 0, and its root port has counted no corrected errors since the boot, where it had counted about one a second for twelve days. A cold power cycle may have retrained a link that trained badly on 18 September. That needs hours of watching before anyone concludes the slot is fine.
- aifoundry2's
/tmpwas cleared by the reboot, as it is at every boot: 19 GB of our working files since 28 September were lost (the repository and home directories are intact).
Why the host hung: candidates, none confirmed
- A hard hang, not a panic. The firmware's panic store was empty at the next boot (
systemd-pstorefound nothing to archive), so the kernel most likely did not panic. (It would not have rebooted anyway:kernel.panicis 0 on these hosts.) - The link went down, and the host tripped over the vanished device. The root port has AER and downstream port containment (DPC) under the kernel's control. If the link failed to train at 8 GT/s, DPC or AER recovery would take card 0 away from under the
et_soc1driver, and an access to it could stall a CPU (on Intel hosts a completion timeout on a memory-mapped read can freeze the machine), or the driver's error path could deadlock. - The platform does not take a forced speed change. The 11th-generation Core CPU's PCIe 4.0 root port may not tolerate a speed change forced with
setpciwhile a driver is bound, whatever the card does. - Card 0 itself. Its link was already marginal (the error flood), and it is the card that overheats; a retrain may have failed on the card's side.
The previous boot's kernel log would narrow this down, but reading it needs root or the adm group: journalctl -b -1 -k -o short-precise | tail -80 as root on aifoundry1.
Lessons
- A link-level experiment is a host-level risk. Treat retraining, re-speeding or resetting a PCIe link like a reboot: only with someone on site, a quiet host, and the users told. Our plan's “only card 0 is at risk” was wrong.
- No out-of-band access means every hang costs a trip. An IP-KVM or a named on-site contact (request SH6) would have turned 26 minutes into two.
- A power cycle clears
/tmpon every machine it touches. Keep work in the home directory. - Plans that change hardware state need a “what if the host goes” line, not only a “what if the device goes” line.
Recorded in the repository: docs/findings/14-card-behaviour.md and the rules table of AGENT.md. Source of this page: docs/reports/2026-09-30-aifoundry1-link-retrain-hang.html.