Analysis: AI regulation based on the text of system prompts is insufficient, say authors writing for Tech Policy Press
An analysis on Tech Policy Press (authors Neumann, Sargeant, Singh) argues that AI safety regulation based on the text of system prompts is insufficient – prompts provide only weak assurances about behavior and can conceal the gap between intent and actual model output. Evidence: incidents involving Grok and GPT-4o.
Companies developing AI models should prepare for the possibility that future regulation may (and, according to the authors, should) require evidence of actual model behavior – adversarial testing, audits across contexts and languages, changelogs – rather than merely submitting the text of a system prompt as proof of safety.