Catalog
affaan-m/data-throughput-accelerator

affaan-m

data-throughput-accelerator

Use when large data ingestion, backfill, export, ETL, warehouse loading, manifest catch-up, or table synchronization needs to become much faster while preserving data correctness.

New~619Updated Jul 14, 2026

Data Throughput Accelerator

Use this skill when the bottleneck is moving, transforming, or saving lots of data. The goal is not just speed. The goal is faster correct data landing in the right place with proof.

First Distinction

Separate these before optimizing:

  • source extraction speed;
  • network transfer speed;
  • warehouse/load speed;
  • transform speed;
  • serving-table freshness;
  • live tail growth while the job runs.

A pipeline can be "fast" and still appear behind if new data arrives faster than the final catch-up window.

Fast Path Heuristics

  • Move compute to where the data already is.
  • Prefer warehouse-native scans, joins, and appends for large landed files.
  • Use manifests or checkpoints so completed files/partitions are skipped.
  • Use partitioning and clustering that match the read and append pattern.
  • Batch small files, requests, and writes.
  • Make writes idempotent through unique keys, manifests, or replaceable staging.
  • Keep raw, derived, and serving tables separately accountable.

Workflow

  1. Read the current source, target, and manifest contracts.
  2. Measure backlog: external files, manifest rows, raw rows, derived rows, min/max timestamps, and unprocessed counts.
  3. Run a safe catch-up or sample benchmark.
  4. Compare variants: batch size, worker count, warehouse SQL, file grouping, staging shape, and manifest update method.
  5. Promote only the fastest path that keeps counts and timestamps coherent.
  6. Codify the path as a CLI, scheduled job, workflow, or runbook.
  7. Rerun final accounting after the codified path executes.

Accounting Output

Use a hard accounting block:

Data throughput result:
- Source files discovered: 294
- Files processed this run: 294
- Raw rows added: 9,683,598
- Derived rows added: 8,917,585
- Remaining tail: 24 files at readback time
- Runtime: 38.7s
- Correctness gate: manifest counts and table max timestamps match

Guardrails

  • Do not delete raw data to make a metric look better.
  • Do not skip failed files silently.
  • Do not mix historical backfill status with live-tail freshness.
  • Do not call a pipeline complete until the target tables and manifest agree.
  • For finance, healthcare, regulated, or customer-impacting data, preserve replay evidence and approval gates.
Files1
1 files · 1.0 KB

Select a file to preview

Overall Score

87/100

Grade

A

Excellent

Safety

88

Quality

86

Clarity

89

Completeness

82

Summary

This skill provides a structured methodology for accelerating data throughput in ETL pipelines, data warehousing, and manifest-driven workflows. It guides agents through source identification, bottleneck analysis, safe benchmarking, and promotion of optimized paths while maintaining data correctness through accounting and guardrails.

Detected Capabilities

file readbash executiondata measurement and analysisquery optimization guidancemanifest/checkpoint managementpipeline benchmarking

Trigger Keywords

Phrases that MCP clients use to match this skill to user intent.

speed up data backfilloptimize etl pipelinefaster warehouse loadingaccelerate data exporttable synchronizationmanifest catch-up

Use Cases

  • Accelerating large-scale data backfills into data warehouses
  • Optimizing ETL pipelines to reduce extract-transform-load latency
  • Synchronizing table state across distributed systems using manifests
  • Implementing faster data export workflows without data loss
  • Catching up on live-tail growth during ongoing data ingestion
  • Benchmarking warehouse-native queries to replace slow client-side transforms

Quality Notes

  • Excellent structural organization with clear workflow steps and decision points
  • Strong guardrails section that explicitly forbids common data integrity pitfalls (silent failures, data deletion for metrics, mixing historical/live state)
  • Practical heuristics are well-reasoned and emphasize data-aware optimization (co-locate compute, use warehouse-native operations, batch intelligently)
  • Accounting output template makes data correctness verifiable and auditable
  • Clear distinction of bottleneck categories helps agents diagnose root causes before optimizing
  • Well-suited for regulated/financial domains with emphasis on replay evidence and approval gates
  • Terminology is precise (raw/derived/serving tables, manifest contracts, staging idempotency)
  • Good balance between speed optimization and correctness preservation
Model: claude-haiku-4-5-20251001Analyzed: Jul 14, 2026

Reviews

Add this skill to your library to leave a review.

No reviews yet

Be the first to share your experience.

Version History

  1. v1.1

    Content updated

    ✦ AISKILL.md body unchanged; safety grade improved to A.

    2026-07-14

    Latest
  2. v1.0

    2026-05-25

    View This VersionInitial version

Use affaan-m/data-throughput-accelerator in your dev environment

Command Palette

Search for a command to run...