Catalog
affaan-m/gget

gget CLI and Python workflow for quick genomic database queries, sequence lookup, BLAST-style searches, enrichment checks, and reproducible bioinformatics evidence logs.

New~1.2kUpdated Jul 14, 2026

gget

Use this skill when a task needs quick bioinformatics lookup across genomic reference databases with the gget CLI or Python package.

When to Use

  • Finding Ensembl IDs, gene metadata, transcript details, or sequences.
  • Running quick BLAST or BLAT lookups without building a full local pipeline.
  • Fetching reference genome links and annotations from Ensembl.
  • Querying protein structure, pathway, cancer, expression, or disease-association modules through a single interface.
  • Creating a reproducible first-pass evidence log before moving to heavier tools such as Biopython, Snakemake, Nextflow, BLAST+, or database-specific clients.

Use a dedicated workflow instead of gget when the task requires regulated clinical interpretation, high-throughput production pipelines, or fine-grained control over database versions and local indexes.

Installation

Use a clean Python environment.

python -m venv .venv
. .venv/bin/activate
python -m pip install --upgrade pip
python -m pip install --upgrade gget
gget --help

If uv is available:

uv venv
. .venv/bin/activate
uv pip install gget

Before relying on an older environment, upgrade gget and re-check the module docs. The upstream databases queried by gget change over time.

Basic Patterns

CLI shape:

gget <module> [arguments] [options]

Python shape:

import gget

result = gget.search(["BRCA1"], species="human")
print(result)

Common workflow:

  1. Identify the species, assembly, gene ID type, and database needed.
  2. Check the current module documentation for arguments.
  3. Run a small query first.
  4. Save output with an explicit filename and date.
  5. Record module name, version, arguments, and database assumptions.

Common Modules

Use current upstream docs for exact arguments. These modules are common first choices:

  • gget search: find Ensembl IDs from search terms.
  • gget info: retrieve metadata for Ensembl, UniProt, or related IDs.
  • gget seq: fetch nucleotide or amino-acid sequences.
  • gget ref: retrieve reference genome download links.
  • gget blast: run a quick BLAST query.
  • gget blat: locate a sequence against supported genome assemblies.
  • gget muscle: run multiple sequence alignment.
  • gget diamond: run local sequence alignment against reference sequences.
  • gget alphafold and gget pdb: inspect protein-structure references.
  • gget enrichr, gget opentargets, gget archs4, gget bgee, gget cbio, and gget cosmic: explore enrichment, target, expression, cancer, and disease association data.

Do not assume every module supports every Python version or dependency set. Some optional scientific dependencies have narrower version support than the core package.

Quick Examples

Find genes:

gget search -s human brca1 dna repair -o brca1-search.json

Fetch gene metadata:

gget info ENSG00000012048 -o brca1-info.json

Fetch a sequence:

gget seq ENSG00000012048 -o brca1-seq.fa

Run a small BLAST query:

gget blast "MEEPQSDPSVEPPLSQETFSDLWKLLPEN" -l 10 -o blast-results.json

Python example:

import gget

genes = gget.search(["BRCA1", "DNA repair"], species="human")
info = gget.info(["ENSG00000012048"])
sequence = gget.seq("ENSG00000012048")

Reproducibility Log

For scientific outputs, include enough metadata to replay the query.

| Date | gget version | Module | Query | Species/assembly | Output | Notes |
| --- | --- | --- | --- | --- | --- | --- |
| 2026-05-11 | `gget --version` | search | `BRCA1 DNA repair` | human | `brca1-search.json` | Docs checked before run |

Also record:

  • Python version and environment manager.
  • Any optional dependency installed through gget setup.
  • Database-specific identifiers returned by the query.
  • Whether output is JSON, CSV, FASTA, or a DataFrame export.
  • Any failures that were resolved by upgrading gget.

Review Checklist

  • Did you upgrade or verify the installed gget version?
  • Did you check the current upstream module docs before using arguments?
  • Is the species or assembly explicit?
  • Are identifiers preserved exactly, including Ensembl/UniProt prefixes?
  • Is the result labeled as database output rather than clinical interpretation?
  • Is the query reproducible from the saved command or Python snippet?
  • Are optional dependencies installed in an isolated environment?

References

Files1
1 files · 1.0 KB

Select a file to preview

Overall Score

87/100

Grade

A

Excellent

Safety

88

Quality

88

Clarity

86

Completeness

84

Summary

A skill that guides agents through quick genomic database queries using the `gget` CLI and Python package. It covers searching genes, fetching sequences, running BLAST/BLAT lookups, and querying enrichment/disease data across Ensembl, UniProt, and related databases. The skill emphasizes reproducible workflows, version management, and appropriate scope boundaries—when to use gget versus heavier bioinformatics pipelines.

Detected Capabilities

cli-executionpython-package-importfile-writeenvironment-isolationnetwork-requestexternal-database-query

Trigger Keywords

Phrases that MCP clients use to match this skill to user intent.

genomic database lookupensembl gene searchsequence retrievalblast queryreproducible bioinformaticsgget module

Risk Signals

INFO

Network requests to external genomic databases (Ensembl, UniProt, NCBI)

Section: Common Modules, When to Use
INFO

Python virtual environment setup with elevated pip privileges

Section: Installation
INFO

File writes to project-scoped output files (JSON, FASTA, CSV)

Section: Quick Examples, Common Workflow

Referenced Domains

External domains referenced in skill content, detected by static analysis.

doi.orggithub.compachterlab.github.io

Use Cases

  • Search genomic databases for Ensembl IDs and gene metadata without running full local pipelines
  • Quickly fetch nucleotide or protein sequences from reference databases
  • Run BLAST or BLAT queries to locate sequences in supported genomes
  • Query disease associations, protein structures, and expression data across enrichment modules
  • Create a reproducible first-pass evidence log with versioned database queries before moving to production pipelines like Snakemake or Nextflow
  • Prototype bioinformatics workflows in isolated Python environments with explicit version tracking

Quality Notes

  • Clear separation of when to use gget versus heavier pipelines—excellent scope boundaries
  • Installation instructions specify isolated Python environments (.venv) and explicit version upgrade steps—reduces environmental surprises
  • Comprehensive module reference with examples covering search, info, seq, ref, blast, blat, and enrichment modules
  • Reproducibility log template with metadata checklist (version, date, species, arguments, output format) is exemplary for scientific workflows
  • Well-structured review checklist ensures agents verify upstream docs and database assumptions before queries
  • Quick examples cover both CLI and Python usage patterns
  • References to upstream docs and GitHub repository support agents in verifying current module behavior
  • Edge case coverage: explicitly warns that not all modules support all Python versions
  • Good note about optional scientific dependencies with narrower version support
Model: claude-haiku-4-5-20251001Analyzed: Jul 14, 2026

Reviews

Add this skill to your library to leave a review.

No reviews yet

Be the first to share your experience.

Version History

  1. v1.1

    Content updated

    ✦ AINo behavioral changes detected.

    2026-07-14

    Latest
  2. v1.0

    2026-05-15

    View This VersionInitial version

Use affaan-m/gget in your dev environment

Command Palette

Search for a command to run...