Hook up the model's output to a shell. Remember it doesn't have to be engineered to modern SRE standards, it just needs to not crash. Dirty hacks are on the table. The comparison to diaper-changing is fairly illuminating, actually: even human children, let alone adults, figure out how to take care of that kind of thing, and it fades into the background of their activities. I agree that plausible mechanism is lacking in these AI risk conversations, but ops is not the hard part IMO.