Catalog
affaan-m/benchmark-optimization-loop

affaan-m

benchmark-optimization-loop

Convert 'make it faster' requests into a bounded measured optimization loop — baseline first, generate one-hypothesis variants, benchmark each against a correctness gate, and promote the fastest safe variant with reproducible commands. Use when asked to speed something up, try many variants, run recursive optimization, benchmark latency/throughput/cost, or pick the best implementation by repeated measured tests.

NewUpdated Sep 27, 2026

Benchmark Optimization Loop

Use this skill to convert "make it 20x faster" or "try 50 recursive optimizations" into a bounded measured loop that can actually improve a system.

Required Baseline

Do not optimize until these exist:

  • the operation being optimized;
  • the correctness gate that must stay green;
  • the metric: wall time, p95 latency, rows/sec, cost/run, memory, error rate;
  • the current baseline;
  • the search budget: max variants, max time, max spend, max data impact.

If the user asks for an unrealistic target, keep the ambition but make the loop bounded and measurable.

Loop

  1. Measure the baseline.
  2. Identify bottlenecks from evidence.
  3. Generate variants that test one hypothesis each.
  4. Run variants with the same input shape.
  5. Reject variants that fail correctness, safety, or reproducibility.
  6. Promote the fastest safe variant.
  7. Codify the winning path in a script, command, test, config, or doc.
  8. Rerun the baseline and winner to confirm the delta.

Variant Table

Track variants like this:

Variant | Hypothesis | Command | Time | Correct? | Notes
baseline | current path | npm run job | 120s | yes | stable
batch-500 | fewer round trips | npm run job -- --batch 500 | 42s | yes | winner
parallel-8 | more workers | npm run job -- --workers 8 | 31s | no | rate limited

For recursive or hyperparameter work:

  • persist every run to a ledger;
  • compare against the prior accepted winner, not only the previous run;
  • keep a holdout or replay check;
  • stop when improvement is within noise, correctness fails, cost exceeds the budget, or the search starts changing more variables than it can explain.

Use phrases like "best measured safe variant" instead of "global optimum" unless the search space was actually exhaustive.

Promotion Gate

A variant cannot become the new default until:

  • correctness tests pass;
  • the performance delta is repeated or explained;
  • rollback is obvious;
  • the change is encoded in source control or a durable runbook;
  • the final summary includes exact commands and measurements.
Files1
1 files · 1.0 KB

Select a file to preview

Overall Score

78/100

Grade

B

Good

Grades are signals, not a certification. Always review a skill yourself before use.

Safety

85

Quality

75

Clarity

82

Completeness

68

Summary

A structured methodology for converting vague performance-improvement requests ("make it faster") into bounded optimization loops with baseline measurement, hypothesis-driven variants, benchmarking against correctness gates, and reproducible promotion of the fastest safe implementation. The skill provides a loop structure, variant tracking template, guidance for recursive search, and promotion criteria.

Detected Capabilities

baseline measurementvariant generationcommand execution (benchmarking)data logging and ledger trackingperformance metric collectioncorrectness gate validationsource control integration

Trigger Keywords

Phrases that agents use to match this skill to user intent.

make it fasteroptimize latencybenchmark variantsperformance tuninghyperparameter searchcost optimizationthroughput improvement

Use Cases

  • Speed up slow endpoints with measured variants and correctness validation
  • Optimize database queries through hypothesis-driven experimentation
  • Benchmark and select among multiple implementation strategies
  • Conduct hyperparameter tuning with bounded search budgets and reproducibility
  • Evaluate cost, latency, or throughput tradeoffs in production systems

Quality Notes

  • Clear structure with step-by-step loop that prevents premature optimization and ensures baselines
  • Variant table template is concrete and immediately usable
  • Recursive search guidance explicitly addresses common pitfalls (comparing against prior winner, not just previous run; holding out test data)
  • Promotion gate criteria are practical and enforceable
  • Required Baseline section sets appropriate guardrails and prevents scope creep (e.g., 'unrealistic target' handling)
  • No examples of actual commands or tool invocations — would benefit from concrete bash/npm examples showing how to instrument measurements
  • Limited guidance on how to generate 'one-hypothesis' variants — relies on user creativity
  • No error handling or timeout guidance for long-running benchmarks
  • Correctness gate is assumed to exist but not deeply explained — brief guidance on how to structure or detect correctness failures would strengthen the skill
Model: claude-haiku-4-5-20251001Analyzed: Sep 27, 2026

Reviews

Add this skill to your library to leave a review.

No reviews yet

Be the first to share your experience.

Version History

  1. v2.0

    Contract changed: description

    ✦ AIDescription now emphasizes baseline-first, bounded optimization loop with correctness gates and reproducible commands.

    triggering2026-09-27

    LATEST
  2. v1.2

    Content updated

    ✦ AIAdds MIT license to skill frontmatter.

    license2026-09-09

    View This Version
  3. v1.1

    Content updated

    ✦ AISKILL.md body unchanged; safety grade moved A to B.

    2026-07-14

    View This Version
  4. v1.0

    2026-05-25

    View This VersionInitial version

Use affaan-m/benchmark-optimization-loop in your dev environment

Command Palette

Search for a command to run...