Fighting Fire with Fire: On the Feasibility of Protecting Exercises Against AI Cheating

Protected visual questions steer AI assistants toward controlled wrong answers that form a detectable assignment-level fingerprint.

Assignment fingerprint
Protected questions steer AI solvers toward secret wrong answers. Repeated target matches across an assignment form a detectable fingerprint of sustained blind copying.

01 / Key message

Instead of guessing whether an answer was AI-written, design a set of questions whose controlled wrong-answer pattern reveals sustained blind copying.

02 / Method

How it works

01

Choose a target

Assign each multimodal question a secret, incorrect target answer.

02

Steer the solver

Optimize a subtle visual perturbation against an ensemble of accessible surrogate models.

03

Calibrate transfer

Query frontier assistants repeatedly and retain question-model pairs with reliable target separation.

04

Detect the pattern

Combine retained questions into an assignment and test for unusual overlap with the hidden targets.

03 / Abstract

Abstract

As multimodal assistants solve more educational exercises, detecting copied answers after submission becomes increasingly unreliable. This work explores a preventive alternative: add subtle, task-preserving perturbations to the visual parts of multiple-choice questions so AI solvers are steered toward designated incorrect answers. Across an assignment, those controlled errors form a statistical fingerprint. The method optimizes against accessible surrogate models, calibrates transfer through repeated black-box queries, and assembles questions whose answer patterns support likelihood-ratio testing across Claude, Gemini, and GPT assistants. The study establishes feasibility under a defined sustained-copying threat model while making educator judgment and the method’s limitations explicit.

04 / Contributions

What this adds

  1. 01

    Assessment-side intervention

    Moves the problem from post-hoc authorship classification to proactive exercise design.

  2. 02

    Controlled error fingerprint

    Uses target-specific visual perturbations to create a pattern that blind AI copying reproduces across an assignment.

  3. 03

    Calibrated detection

    Combines black-box response calibration with assistant-specific statistical tests and clearly stated student-model assumptions.

05 / Citation

Citation

Braun, T., Grebe, J., Rethfeld, L., & Rohrbach, M. (2026). Fighting Fire with Fire: On the Feasibility of Protecting Exercises Against AI Cheating. Preprint.

BibTeX
@misc{braun2026fighting,
  title  = {Fighting Fire with Fire: On the Feasibility of Protecting Exercises Against AI Cheating},
  author = {Tobias Braun and Jonas Grebe and Louis Rethfeld and Marcus Rohrbach},
  year   = {2026},
  note   = {Preprint}
}