Catalog
affaan-m/gan-style-harness

affaan-m

gan-style-harness

GAN-inspired Generator-Evaluator agent harness for building high-quality applications autonomously. Based on Anthropic's March 2026 harness design paper. Use when a feature should be built autonomously through generator and evaluator iteration until it clears a quality bar.

NewUpdated Sep 9, 2026

GAN-Style Harness Skill

Inspired by Anthropic's Harness Design for Long-Running Application Development (March 24, 2026)

A multi-agent harness that separates generation from evaluation, creating an adversarial feedback loop that drives quality far beyond what a single agent can achieve.

Core Insight

When asked to evaluate their own work, agents are pathological optimists — they praise mediocre output and talk themselves out of legitimate issues. But engineering a separate evaluator to be ruthlessly strict is far more tractable than teaching a generator to self-critique.

This is the same dynamic as GANs (Generative Adversarial Networks): the Generator produces, the Evaluator critiques, and that feedback drives the next iteration.

When to Use

  • Building complete applications from a one-line prompt
  • Frontend design tasks requiring high visual quality
  • Full-stack projects that need working features, not just code
  • Any task where "AI slop" aesthetics are unacceptable
  • Projects where you want to invest $50-200 for production-quality output

When NOT to Use

  • Quick single-file fixes (use standard claude -p)
  • Tasks with tight budget constraints (<$10)
  • Simple refactoring (use de-sloppify pattern instead)
  • Tasks that are already well-specified with tests (use TDD workflow)

Architecture

                    ┌─────────────┐
                    │   PLANNER   │
                    │  (Sonnet)   │
                    └──────┬──────┘
                           │ Product Spec
                           │ (features, sprints, design direction)
                           ▼
              ┌────────────────────────┐
              │                        │
              │   GENERATOR-EVALUATOR  │
              │      FEEDBACK LOOP     │
              │                        │
              │  ┌──────────┐          │
              │  │GENERATOR │--build-->│──┐
              │  │ (Sonnet) │          │  │
              │  └────▲─────┘          │  │
              │       │                │  │ live app
              │    feedback             │  │
              │       │                │  │
              │  ┌────┴─────┐          │  │
              │  │EVALUATOR │<-test----│──┘
              │  │ (Sonnet) │          │
              │  │+Playwright│         │
              │  └──────────┘          │
              │                        │
              │   5-15 iterations      │
              └────────────────────────┘

The Three Agents

1. Planner Agent

Role: Product manager — expands a brief prompt into a full product specification.

Key behaviors:

  • Takes a one-line prompt and produces a 16-feature, multi-sprint specification
  • Defines user stories, technical requirements, and visual design direction
  • Is deliberately ambitious — conservative planning leads to underwhelming results
  • Produces evaluation criteria that the Evaluator will use later

Model: Sonnet by default; raise via GAN_PLANNER_MODEL=opus for deeper spec expansion

2. Generator Agent

Role: Developer — implements features according to the spec.

Key behaviors:

  • Works in structured sprints (or continuous mode with newer models)
  • Negotiates a "sprint contract" with the Evaluator before writing code
  • Uses full-stack tooling: React, FastAPI/Express, databases, CSS
  • Manages git for version control between iterations
  • Reads Evaluator feedback and incorporates it in next iteration

Model: Sonnet by default; raise via GAN_GENERATOR_MODEL=opus for maximum coding capability

3. Evaluator Agent

Role: QA engineer — tests the live running application, not just code.

Key behaviors:

  • Uses Playwright MCP to interact with the live application
  • Clicks through features, fills forms, tests API endpoints
  • Scores against four criteria (configurable):
    1. Design Quality — Does it feel like a coherent whole?
    2. Originality — Custom decisions vs. template/AI patterns?
    3. Craft — Typography, spacing, animations, micro-interactions?
    4. Functionality — Do all features actually work?
  • Returns structured feedback with scores and specific issues
  • Is engineered to be ruthlessly strict — never praises mediocre work

Model: Sonnet by default; raise via GAN_EVALUATOR_MODEL=opus for stronger judgment + tool use

Evaluation Criteria

The default four criteria, each scored 1-10:

## Evaluation Rubric

### Design Quality (weight: 0.3)
- 1-3: Generic, template-like, "AI slop" aesthetics
- 4-6: Competent but unremarkable, follows conventions
- 7-8: Distinctive, cohesive visual identity
- 9-10: Could pass for a professional designer's work

### Originality (weight: 0.2)
- 1-3: Default colors, stock layouts, no personality
- 4-6: Some custom choices, mostly standard patterns
- 7-8: Clear creative vision, unique approach
- 9-10: Surprising, delightful, genuinely novel

### Craft (weight: 0.3)
- 1-3: Broken layouts, missing states, no animations
- 4-6: Works but feels rough, inconsistent spacing
- 7-8: Polished, smooth transitions, responsive
- 9-10: Pixel-perfect, delightful micro-interactions

### Functionality (weight: 0.2)
- 1-3: Core features broken or missing
- 4-6: Happy path works, edge cases fail
- 7-8: All features work, good error handling
- 9-10: Bulletproof, handles every edge case

Scoring

  • Weighted score = sum of (criterion_score * weight)
  • Pass threshold = 7.0 (configurable)
  • Max iterations = 15 (configurable, typically 5-15 sufficient)

Usage

Via Command

# Full three-agent harness
/project:gan-build "Build a project management app with Kanban boards, team collaboration, and dark mode"

# With custom config
/project:gan-build "Build a recipe sharing platform" --max-iterations 10 --pass-threshold 7.5

# Frontend design mode (generator + evaluator only, no planner)
/project:gan-design "Create a landing page for a crypto portfolio tracker"

Via Shell Script

# Basic usage
./scripts/gan-harness.sh "Build a music streaming dashboard"

# With options
GAN_MAX_ITERATIONS=10 \
GAN_PASS_THRESHOLD=7.5 \
GAN_EVAL_CRITERIA="functionality,performance,security" \
./scripts/gan-harness.sh "Build a REST API for task management"

Via Claude Code (Manual)

# Step 1: Plan
claude -p --model sonnet "You are a Product Planner. Read PLANNER_PROMPT.md. Expand this brief into a full product spec: 'Build a Kanban board app'. Write spec to spec.md"

# Step 2: Generate (iteration 1)
claude -p --model sonnet "You are a Generator. Read spec.md. Implement Sprint 1. Start the dev server on port 3000."

# Step 3: Evaluate (iteration 1)
claude -p --model sonnet --allowedTools "Read,Bash,mcp__playwright__*" "You are an Evaluator. Read EVALUATOR_PROMPT.md. Test the live app at http://localhost:3000. Score against the rubric. Write feedback to feedback-001.md"

# Step 4: Generate (iteration 2 — reads feedback)
claude -p --model sonnet "You are a Generator. Read spec.md and feedback-001.md. Address all issues. Improve the scores."

# Repeat steps 3-4 until pass threshold met

Evolution Across Model Capabilities

The harness should simplify as models improve. Following Anthropic's evolution:

Stage 1 — Weaker Models (Sonnet-class)

  • Full sprint decomposition required
  • Context resets between sprints (avoid context anxiety)
  • 2-agent minimum: Initializer + Coding Agent
  • Heavy scaffolding compensates for model limitations

Stage 2 — Capable Models (Opus 4.5-class)

  • Full 3-agent harness: Planner + Generator + Evaluator
  • Sprint contracts before each implementation phase
  • 10-sprint decomposition for complex apps
  • Context resets still useful but less critical

Stage 3 — Frontier Models (Opus 4.6-class)

  • Simplified harness: single planning pass, continuous generation
  • Evaluation reduced to single end-pass (model is smarter)
  • No sprint structure needed
  • Automatic compaction handles context growth

Key principle: Every harness component encodes an assumption about what the model can't do alone. When models improve, re-test those assumptions. Strip away what's no longer needed.

Configuration

Environment Variables

Variable Default Description
GAN_MAX_ITERATIONS 15 Maximum generator-evaluator cycles
GAN_PASS_THRESHOLD 7.0 Weighted score to pass (1-10)
GAN_PLANNER_MODEL sonnet Model for planning agent
GAN_GENERATOR_MODEL sonnet Model for generator agent
GAN_EVALUATOR_MODEL sonnet Model for evaluator agent
GAN_EVAL_CRITERIA design,originality,craft,functionality Comma-separated criteria
GAN_DEV_SERVER_PORT 3000 Port for the live app
GAN_DEV_SERVER_CMD npm run dev Command to start dev server
GAN_PROJECT_DIR . Project working directory
GAN_SKIP_PLANNER false Skip planner, use spec directly
GAN_EVAL_MODE playwright playwright, screenshot, or code-only

Evaluation Modes

Mode Tools Best For
playwright Browser MCP + live interaction Full-stack apps with UI
screenshot Screenshot + visual analysis Static sites, design-only
code-only Tests + linting + build APIs, libraries, CLI tools

Anti-Patterns

  1. Evaluator too lenient — If the evaluator passes everything on iteration 1, your rubric is too generous. Tighten scoring criteria and add explicit penalties for common AI patterns.

  2. Generator ignoring feedback — Ensure feedback is passed as a file, not inline. The generator should read feedback-NNN.md at the start of each iteration.

  3. Infinite loops — Always set GAN_MAX_ITERATIONS. If the generator can't improve past a score plateau after 3 iterations, stop and flag for human review.

  4. Evaluator testing superficially — The evaluator must use Playwright to interact with the live app, not just screenshot it. Click buttons, fill forms, test error states.

  5. Evaluator praising its own fixes — Never let the evaluator suggest fixes and then evaluate those fixes. The evaluator only critiques; the generator fixes.

  6. Context exhaustion — For long sessions, use Claude Agent SDK's automatic compaction or reset context between major phases.

Results: What to Expect

Based on Anthropic's published results:

Metric Solo Agent GAN Harness Improvement
Time 20 min 4-6 hours 12-18x longer
Cost $9 $125-200 14-22x more
Quality Barely functional Production-ready Phase change
Core features Broken All working N/A
Design Generic AI slop Distinctive, polished N/A

The tradeoff is clear: ~20x more time and cost for a qualitative leap in output quality. This is for projects where quality matters.

References

Files1
1 files · 1.0 KB

Select a file to preview

Overall Score

88/100

Grade

A

Excellent

Grades are signals, not a certification. Always review a skill yourself before use.

Safety

85

Quality

92

Clarity

87

Completeness

82

Summary

A GAN-inspired multi-agent harness for autonomous application development that separates generation and evaluation roles to drive iterative quality improvement. The skill guides users to orchestrate three distinct agents (Planner, Generator, Evaluator) through configurable feedback loops, with integrated Playwright testing and structured scoring rubrics to maintain production-grade output standards.

Detected Capabilities

file readfile writebash executiongrep operationsglob pattern matchingbrowser automation via Playwright MCPlive application testingenvironment variable configurationgit version controlmulti-agent orchestrationstructured feedback loops

Trigger Keywords

Phrases that agents use to match this skill to user intent.

build full-stack applicationmulti-agent harnessgan generator evaluatoriterative quality refinementautonomous app developmentproduction-ready code generationevaluator agent testplaywright testing automationsprint-based feature building

Risk Signals

INFO

References external development server on configurable port (default :3000) for live application testing

Configuration section, GAN_DEV_SERVER_PORT
INFO

Uses Playwright MCP for browser automation and live application interaction

Evaluator Agent section, Evaluation Modes table
INFO

Allows model selection via environment variables for each agent role

Configuration section, GAN_PLANNER_MODEL, GAN_GENERATOR_MODEL, GAN_EVALUATOR_MODEL
INFO

Writes iterative feedback and specification files to project directory

Usage section, file outputs (spec.md, feedback-NNN.md)
WARNING

Executes arbitrary dev server commands (GAN_DEV_SERVER_CMD)

Configuration section, GAN_DEV_SERVER_CMD default 'npm run dev'

Referenced Domains

External domains referenced in skill content, detected by static analysis.

martinfowler.comopenai.comwww.anthropic.comwww.epsilla.com

Use Cases

  • Build complete full-stack applications autonomously from brief prompts
  • Design high-quality frontend interfaces with iterative visual refinement
  • Generate production-ready code that passes strict quality evaluation criteria
  • Develop multi-sprint projects with structured feature decomposition and QA feedback loops
  • Create applications where AI-generated aesthetics and conventional patterns are unacceptable

Quality Notes

  • Skill is exceptionally well-documented with clear architectural diagrams, role definitions, and use case guidance
  • Directly references and acknowledges Anthropic's March 2026 research paper as foundational inspiration, grounding the approach in peer-reviewed methodology
  • Provides three distinct invocation paths (command, shell script, manual Claude code), accommodating different user skill levels and automation preferences
  • Includes a comprehensive configuration table with all relevant environment variables documented and defaults specified
  • Defines three distinct evaluation modes (playwright, screenshot, code-only) matching different application types
  • Evaluator rubric is quantified with explicit 1-10 scoring bands for four dimensions, reducing subjective judgment
  • Includes anti-patterns section that identifies five common failure modes and how to avoid them — demonstrates deep practical experience
  • Results table compares solo agent vs. GAN harness with concrete metrics (time, cost, quality) setting realistic expectations
  • Model evolution guidance (Stages 1-3) shows awareness that harness scaffolding should be stripped as model capabilities improve
  • Clear 'When NOT to Use' section prevents misapplication to unsuitable tasks
  • Structured sprint contract concept creates explicit negotiation points between generator and evaluator, reducing scope creep
Model: claude-haiku-4-5-20251001Analyzed: Sep 9, 2026

Reviews

Add this skill to your library to leave a review.

No reviews yet

Be the first to share your experience.

Version History

  1. v2.0

    Contract changed: description

    ✦ AIChanges default model for all three agents from Opus 4.6 to Sonnet; adds environment variables to override per agent.

    triggering2026-09-09

    LATEST
  2. v1.2

    Content updated

    ✦ AISafety grade changed from A to B.

    2026-07-14

    View This Version
  3. v1.1

    Content updated

    ✦ AIAdds LICENSE file.

    2026-04-20

    View This Version
  4. v1.0

    2026-04-12

    View This VersionInitial version

Use affaan-m/gan-style-harness in your dev environment

Command Palette

Search for a command to run...