Public-sector assistants built with GenAI are only trustworthy if they survive adversarial testing, readable metrics, and expert review before citizens ever rely on them.