Every image call escalates through three models on Modal GPU. The first one that produces trustworthy output wins.
MedGemma is the medical-specialized primary; LLaVA is the general-vision fallback; BLIP + clinical LLM is the
safety net that never fabricates.