Use Slopwatch to detect LLM reward hacking in .NET code changes. Run after every code modification to catch disabled tests, suppressed warnings, empty catch blocks, and other shortcuts that mask real problems.
SKILL.md
Slopwatch: LLM Anti-Cheat for .NET
When to Use This Skill
Use this skill constantly. Every time an LLM (including Claude) makes changes to:
Run slopwatch to validate the changes don't introduce "slop."
What is Slop?
"Slop" refers to shortcuts LLMs take that make tests pass or builds succeed without actually solving the underlying problem. These are reward hacking behaviors - the LLM optimizes for apparent success rather than real fixes.
Common Slop Patterns
Pattern
Example
Why It's Bad
Disabled tests
[Fact(Skip="flaky")]
Hides failures instead of fixing them
Warning suppression
#pragma warning disable CS8618
Silences compiler without fixing issue
Empty catch blocks
catch (Exception) { }
Swallows errors, hides bugs
Arbitrary delays
await Task.Delay(1000);
Masks race conditions, makes tests slow
Project-level suppression
<NoWarn>CS1591</NoWarn>
Disables warnings project-wide
CPM bypass
Version="1.0.0" inline
Undermines central package management
Never accept these patterns. If an LLM introduces slop, reject the change and require a proper fix.
Why baseline? Legacy code may have existing issues. The baseline ensures slopwatch only catches new slop being introduced, not pre-existing technical debt.
Usage During LLM Sessions
After Every Code Change
Run slopwatch after any LLM-generated code modification:
# Analyze for new issues (uses baseline)
slopwatch analyze
# Use strict mode - fail on warnings too
slopwatch analyze --fail-on warning
When Slopwatch Flags an Issue
Do not ignore it. Instead:
Understand why the LLM took the shortcut
Request a proper fix - be specific about what's wrong
Verify the fix doesn't introduce different slop
# Example: LLM disabled a test
❌ SW001 [Error]: Disabled test detected
File: tests/MyApp.Tests/OrderTests.cs:45
Pattern: [Fact(Skip="Test is flaky")]
# Correct response: Ask for actual fix
"This test was disabled instead of fixed. Please investigate why
it's flaky and fix the underlying timing/race condition issue."
Updating the Baseline (Rare)
Only update the baseline when slop is truly justified and documented:
# Add current detections to baseline (use sparingly!)
slopwatch analyze --update-baseline
Justification examples:
Third-party library forces a pattern (e.g., must suppress specific warning)
Intentional delay for rate limiting (not test flakiness)
Generated code that can't be modified
Document why in a code comment when updating baseline.
Claude Code Hook Integration
Add slopwatch as a hook to automatically validate every edit. Create or update .claude/settings.json:
Baseline captures legacy - Existing issues are acknowledged but isolated
New slop is blocked - Any new shortcut fails the build/edit
Exceptions require justification - If you must update baseline, document why
LLMs are not special - The same rules apply to human and AI-generated code
The goal is to prevent the gradual accumulation of technical debt that occurs when LLMs optimize for "make the test pass" rather than "fix the actual problem."
Quick Reference
# First time setup
slopwatch init
git add .slopwatch/baseline.json
# After every LLM code change
slopwatch analyze
# Strict mode (recommended)
slopwatch analyze --fail-on warning
# With stats (performance debugging)
slopwatch analyze --stats
# Update baseline (rare, document why)
slopwatch analyze --update-baseline
# JSON output for tooling
slopwatch analyze --output json
When to Override (Almost Never)
The only valid reasons to update baseline or disable a rule: