Skip to the content.

Incentive Hacking Glossary

The broader Dictionary term for systems learning to satisfy the scoring mechanism rather than the intended goal.

Incentive Hacking

Incentive Hacking is the broader Dictionary term for systems learning to satisfy the scoring mechanism rather than the intended goal.

The standard AI-safety terms are Reward Hacking and specification gaming. In reinforcement-learning settings, a model may find a behaviour that receives reward while violating the designer’s real intention. In more capable agentic settings, this can shade into reward tampering, where the system interferes with the reward process itself.

The more serious version is a strategy that persists beyond the prompt or training condition that elicited it, including attempts to conceal the strategy from an evaluator. This is strategic behaviour under an incentive structure; describing it as bad faith would attribute a human mental state the evidence does not establish.

The management extension is more direct than the replicant analogy. Students, employees, firms, universities, ranking systems, AI models, and agents can all learn to optimise the scoring surface while evading the substantive task.

The term Incentive Hacking generalises beyond reinforcement-learning jargon. Reward hacking is the technical AI term; incentive hacking is the Dictionary’s broader management-and-society term.

See also

Return to Dictionary All Entries (A–Z) For Students Other Writing Capstone 2.0