Catalog
affaan-m/data-throughput-accelerator

affaan-m

data-throughput-accelerator

Diagnose and accelerate large data movement — ingestion, backfill, export, ETL, warehouse loading, manifest catch-up, and table synchronization — by isolating the true bottleneck, benchmarking variants, and codifying the fastest path with a hard accounting block proving rows and timestamps cohere. Use when a pipeline or backfill is too slow and must get faster without losing data correctness.

NewUpdated Sep 27, 2026

Data Throughput Accelerator

Use this skill when the bottleneck is moving, transforming, or saving lots of data. The goal is not just speed. The goal is faster correct data landing in the right place with proof.

First Distinction

Separate these before optimizing:

  • source extraction speed;
  • network transfer speed;
  • warehouse/load speed;
  • transform speed;
  • serving-table freshness;
  • live tail growth while the job runs.

A pipeline can be "fast" and still appear behind if new data arrives faster than the final catch-up window.

Fast Path Heuristics

  • Move compute to where the data already is.
  • Prefer warehouse-native scans, joins, and appends for large landed files.
  • Use manifests or checkpoints so completed files/partitions are skipped.
  • Use partitioning and clustering that match the read and append pattern.
  • Batch small files, requests, and writes.
  • Make writes idempotent through unique keys, manifests, or replaceable staging.
  • Keep raw, derived, and serving tables separately accountable.

Workflow

  1. Read the current source, target, and manifest contracts.
  2. Measure backlog: external files, manifest rows, raw rows, derived rows, min/max timestamps, and unprocessed counts.
  3. Run a safe catch-up or sample benchmark.
  4. Compare variants: batch size, worker count, warehouse SQL, file grouping, staging shape, and manifest update method.
  5. Promote only the fastest path that keeps counts and timestamps coherent.
  6. Codify the path as a CLI, scheduled job, workflow, or runbook.
  7. Rerun final accounting after the codified path executes.

Accounting Output

Use a hard accounting block:

Data throughput result:
- Source files discovered: 294
- Files processed this run: 294
- Raw rows added: 9,683,598
- Derived rows added: 8,917,585
- Remaining tail: 24 files at readback time
- Runtime: 38.7s
- Correctness gate: manifest counts and table max timestamps match

Guardrails

  • Do not delete raw data to make a metric look better.
  • Do not skip failed files silently.
  • Do not mix historical backfill status with live-tail freshness.
  • Do not call a pipeline complete until the target tables and manifest agree.
  • For finance, healthcare, regulated, or customer-impacting data, preserve replay evidence and approval gates.
Files1
1 files · 1.0 KB

Select a file to preview

Overall Score

76/100

Grade

B

Good

Grades are signals, not a certification. Always review a skill yourself before use.

Safety

92

Quality

72

Clarity

78

Completeness

62

Summary

A data pipeline optimization guide that teaches developers to isolate bottlenecks in large-scale data movement (ingestion, backfill, ETL, warehouse loading). It emphasizes a systematic workflow: measure backlog, benchmark variants, compare performance, and validate correctness through hard accounting blocks that prove row counts and timestamps cohere.

Detected Capabilities

file readingdata measurement and backlog analysisbenchmarking and variant comparisonSQL execution against warehousesmanifest and checkpoint managementaccounting and correctness validationCLI or scheduled job codification

Trigger Keywords

Phrases that agents use to match this skill to user intent.

pipeline too slowslow backfillaccelerate data loadetl bottleneckdata ingestion speedwarehouse throughputcatch-up manifestbatch optimization

Risk Signals

INFO

No destructive operations or credential access patterns detected

Full content scan
INFO

Guardrails section explicitly prohibits dangerous practices (silent failures, deleting raw data, skipping failed files)

Guardrails section
INFO

Emphasis on preservation of replay evidence and approval gates for regulated data

Guardrails section

Use Cases

  • Accelerating slow data pipelines and backfills without data loss
  • Identifying bottlenecks in extraction, transfer, transformation, and load stages
  • Benchmarking batch sizes, worker counts, and SQL query variants for optimal throughput
  • Validating data correctness through manifest and table timestamp reconciliation
  • Designing idempotent writes and skip logic for resumed or replayed loads
  • Optimizing warehouse-native operations for large landed files

Quality Notes

  • Skill is domain-specific and focused on a clear problem: throughput acceleration with correctness validation
  • Workflow is concrete and sequenced (7 steps from measurement to codification)
  • Guardrails section directly addresses common data pipeline anti-patterns (silent failures, data deletion, mixing backfill and live-tail metrics)
  • Heuristics are actionable (move compute to data, use warehouse-native operations, batch small requests, make writes idempotent)
  • Accounting output example is precise and includes key metrics (file counts, row counts, min/max timestamps, correctness gate)
  • Strong emphasis on data integrity as a first-class requirement alongside speed
  • Lacks concrete code examples or SQL patterns — relies on practitioner knowledge of their specific pipeline tool (Airflow, dbt, Spark, etc.)
  • Does not provide templates for manifest structures, checkpoint formats, or variant comparison tables
  • Limited guidance on how to measure 'correct' timestamps when data arrives out-of-order or during backfill
  • No reference to specific tools (dbt, Airflow, Fivetran, Snowflake, BigQuery) — skill assumes agent adapts workflow to their stack
Model: claude-haiku-4-5-20251001Analyzed: Sep 27, 2026

Reviews

Add this skill to your library to leave a review.

No reviews yet

Be the first to share your experience.

Version History

  1. v2.0

    Contract changed: description

    ✦ AISkill now activates when diagnosing bottlenecks, benchmarking variants, and validating data coherence in data movement pipelines—not just speed optimization.

    triggering2026-09-27

    LATEST
  2. v1.2

    Content updated

    ✦ AIAdds MIT license declaration.

    license2026-09-09

    View This Version
  3. v1.1

    Content updated

    ✦ AISKILL.md body unchanged; safety grade improved to A.

    2026-07-14

    View This Version
  4. v1.0

    2026-05-25

    View This VersionInitial version

Use affaan-m/data-throughput-accelerator in your dev environment

Command Palette

Search for a command to run...