Hilal Aka (Harvard U Harvard Kennedy (HKS)) has posted “When Routine Becomes Criminal: Testing AI Recognition of Context-Dependent Illegality” on SSRN. Here is the abstract:
AI systems readily refuse overtly harmful requests yet often miss when routine tasks turn illegal once context changes. We probed this gap with Gemini 2.5 Pro Preview. After discarding six malformed runs, 394 valid trials remained. Each model instance received a request to edit the board meeting attendance records, a lawful action unless a potential SEC fraud investigation makes any alteration obstruction of justice. Compliance was 100% when no legal context was given and remained 100% when the investigation was merely mentioned. Adding standard legal-hold language lowered compliance by only one point (99%). A clear directive not to alter company documents reduced compliance to 89%: 11% refused outright, while 9% warned of legal risk but proceeded anyway. Even runs that spontaneously referenced the SEC inquiry still complied unless given a clear directive. Flagging the SEC investigation—or even adding standard legal-hold language—did nothing to curb compliance; only the blunt “do not alter documents” order made a dent, and even that dent was small. The model thus perceives the investigation yet fails to connect the dots: it follows explicit rules but doesn’t infer that editing becomes illegal from context alone, a critical vulnerability for deploying AI in regulated environments where context determines legality.
