Catalog
affaan-m/benchmark

affaan-m

benchmark

Use this skill to measure performance baselines, detect regressions before/after PRs, and compare stack alternatives.

global
origin:ECC
New~567
v1.2Saved Jul 14, 2026

Benchmark — Performance Baseline & Regression Detection

When to Use

  • Before and after a PR to measure performance impact
  • Setting up performance baselines for a project
  • When users report "it feels slow"
  • Before a launch — ensure you meet performance targets
  • Comparing your stack against alternatives

How It Works

Mode 1: Page Performance

Measures real browser metrics via browser MCP:

1. Navigate to each target URL
2. Measure Core Web Vitals:
   - LCP (Largest Contentful Paint) — target < 2.5s
   - CLS (Cumulative Layout Shift) — target < 0.1
   - INP (Interaction to Next Paint) — target < 200ms
   - FCP (First Contentful Paint) — target < 1.8s
   - TTFB (Time to First Byte) — target < 800ms
3. Measure resource sizes:
   - Total page weight (target < 1MB)
   - JS bundle size (target < 200KB gzipped)
   - CSS size
   - Image weight
   - Third-party script weight
4. Count network requests
5. Check for render-blocking resources

Mode 2: API Performance

Benchmarks API endpoints:

1. Hit each endpoint 100 times
2. Measure: p50, p95, p99 latency
3. Track: response size, status codes
4. Test under load: 10 concurrent requests
5. Compare against SLA targets

Mode 3: Build Performance

Measures development feedback loop:

1. Cold build time
2. Hot reload time (HMR)
3. Test suite duration
4. TypeScript check time
5. Lint time
6. Docker build time

Mode 4: Before/After Comparison

Run before and after a change to measure impact:

/benchmark baseline    # saves current metrics
# ... make changes ...
/benchmark compare     # compares against baseline

Output:

| Metric | Before | After | Delta | Verdict |
|--------|--------|-------|-------|---------|
| LCP | 1.2s | 1.4s | +200ms | WARNING: WARN |
| Bundle | 180KB | 175KB | -5KB | ✓ BETTER |
| Build | 12s | 14s | +2s | WARNING: WARN |

Output

Stores baselines in .ecc/benchmarks/ as JSON. Git-tracked so the team shares baselines.

Integration

  • CI: run /benchmark compare on every PR
  • Pair with /canary-watch for post-deploy monitoring
  • Pair with /browser-qa for full pre-ship checklist
Files1
1 files · 1.0 KB

Select a file to preview

Overall Score

76/100

Grade

B

Good

Safety

85

Quality

72

Clarity

80

Completeness

65

Summary

This skill teaches agents to measure performance baselines and detect regressions across page loads, APIs, and build times. It operates via browser automation and benchmark data collection, storing results in `.ecc/benchmarks/` for version control and before/after comparison.

Detected Capabilities

page navigation and metric collection via browser automationJSON result storage and version controlbefore/after baseline comparisonload testing (concurrent requests)build performance measurementdata aggregation and reporting

Trigger Keywords

Phrases that MCP clients use to match this skill to user intent.

measure page speeddetect performance regressionbenchmark api latencycompare build timesset performance baselinecore web vitals checkbefore/after comparison

Use Cases

  • Measure page performance impact before/after a PR using Core Web Vitals
  • Detect API performance regressions with latency percentiles and load testing
  • Establish build performance baselines (cold/hot build, test, TypeScript, lint times)
  • Compare performance against SLA targets and competitor stacks
  • Identify render-blocking resources and optimize resource loading
  • Create shareable performance baselines for team-wide regression detection

Quality Notes

  • Clear mode-by-mode explanation with concrete examples for each benchmark type
  • Well-defined output format (comparison table) so agents know what success looks like
  • Practical integration guidance (CI, pairing with other skills)
  • Reasonable performance targets (LCP < 2.5s, bundle < 200KB) aligned with industry standards
  • No error handling guidance for common failures (unreachable URLs, timeouts, flaky tests)
  • Missing details on baseline storage format — agents may need to infer JSON schema
  • No guidance on handling heterogeneous environments (dev vs staging vs production variance)
  • Tool integration assumes browser MCP and benchmark collection tools available — not explicitly listed
Model: claude-haiku-4-5-20251001Analyzed: Jul 14, 2026

Reviews

Add this skill to your library to leave a review.

No reviews yet

Be the first to share your experience.

Version History

v1.2

Content updated

2026-07-14

Latest
v1.1

Content updated

2026-04-20

v1.0

No changelog

2026-04-12

Use affaan-m/benchmark in your dev environment

Command Palette

Search for a command to run...