Connection check
verified live · 20h ago
genomic-intelligence
Hosted DNA language models: promoter, splice, enhancer, chromatin, expression, annotation
Tools
15
GitHub stars
—
Installs / wk
—
Licence
—
Transport
streamable-http
Last checked
20h ago
Tools & capabilities
15 toolsRead from the running server on 20h ago.
fetch_ensembl_sequence
gene*speciesflank_bp
Fetch a gene's reference sequence from Ensembl and store it. Returns a handle ({ref, name, length, preview, ...}). Pass the `ref` to predict_* tools — the bases st… Fetch a gene's reference sequence from Ensembl and store it. Returns a handle ({ref, name, length, preview, ...}). Pass the `ref` to predict_* tools — the bases stay server-side. For expression, use fetch_gene_for_expression instead (it prepares the TSS-centred window that model needs).
fetch_gene_for_expression
gene*species
Fetch a gene's sequence prepared for expression prediction. Resolves the gene's TSS via Ensembl and returns the exact TSS-centred 9,198 bp window the expression mo… Fetch a gene's sequence prepared for expression prediction. Resolves the gene's TSS via Ensembl and returns the exact TSS-centred 9,198 bp window the expression model scores, as a handle to pass to predict_expression(sequence_ref=...). Because the window is exactly 9,198 bp, no `tss_index` is needed on that call.
fetch_region
region*strandspeciesflank_bp
Fetch a genomic region by coordinates from Ensembl and store it. For "find the genes in chr8:127,680,000-127,800,000"-style requests: resolves a coordinate range t… Fetch a genomic region by coordinates from Ensembl and store it. For "find the genes in chr8:127,680,000-127,800,000"-style requests: resolves a coordinate range to reference sequence and returns a handle ({ref, name, length, ...}) to pass to find_genes / predict_* — the bases stay server-side. Plus strand by default, which is what the gene-finder expects. For a gene by name use fetch_ensembl_sequence; for expression use fetch_gene_for_expression.
find_genes
read-only
waitmodelsequencesequence_refsequence_name
Find genes (transcript intervals) in a genomic region (async, ~8-25s). Takes 1,000–500,000 bp. The floor is the strictest of the scanning tasks: gene finding needs… Find genes (transcript intervals) in a genomic region (async, ~8-25s). Takes 1,000–500,000 bp. The floor is the strictest of the scanning tasks: gene finding needs a region, not a site. (Only expression's 9,198 bp is higher, and that is a fixed window rather than a minimum region size.) Gene-finding: detects transcript boundaries (TSS + PolyA) and returns one interval per predicted transcript — start/end, strand, a confidence score, and predicted TSS/PolyA positions (BED-style feature intervals, not free-text notes). Use this for "what genes are here", "find / locate genes", or "annotate this region". Each transcript also carries its type (mRNA/lnc_RNA) and internal exon/intron/CDS structure in `exons`/`introns`/`cds` arrays, plus a browser-ready GFF3 track in `data.formats.gff3`. To get each gene's *expression* from a raw region, use find_genes_and_predict_expression instead — expression needs a per-gene TSS window, so predict_expression cannot run on a whole region. Submits an async job internally. With wait=True (default), blocks and streams progress, then returns the result {data, meta} — it never returns a job_id on this path. (If a generous block ceiling is exceeded it returns a timeout error, not a job handle.) With wait=False (detached), returns {data: {job_id, status: 'submitted'}} immediately — poll it with get_job.
find_genes_and_predict_expression
read-only
waitsequencedescriptionsequence_refsequence_name
Find genes in a sequence, then predict each gene's expression (composite). Server-side chaining in ONE call: finds genes (transcript intervals, with their TSS) in… Find genes in a sequence, then predict each gene's expression (composite). Server-side chaining in ONE call: finds genes (transcript intervals, with their TSS) in the sequence, then predicts expression off each discovered TSS in the given experimental context. This is the right tool whenever you want expression for a raw region or sequence — e.g. "find the genes in chr8:… and predict their expression in K562". predict_expression scores ONE TSS window and needs you to know where that TSS is (either a pre-centred 9,198 bp window or a `tss_index`); this tool discovers every gene's TSS itself. It has no 9,198 bp floor and no tss_index; it starts with gene finding, so it takes 1,000–500,000 bp. Runs async internally at every size (the annotate stage is slow even for small inputs), so progress always streams. With wait=True (default), blocks and streams progress, then returns the result {data, meta} — it never returns a job_id on this path. With wait=False (detached), returns {data: {job_id, status: 'submitted'}} immediately — poll it with get_job. Because it ends in expression, `description` (cell type / assay context) is REQUIRED.
get_job
read-only
job_id*
Poll an async job once. Returns the {data, meta} result if complete, a progress envelope if still running, or an error envelope if it failed. Poll an async job once. Returns the {data, meta} result if complete, a progress envelope if still running, or an error envelope if it failed.
list_jobs
read-only
limit
List the caller's recent async jobs (also available as gi://jobs/recent). List the caller's recent async jobs (also available as gi://jobs/recent).
list_models
read-only
task*
List available models for a task. Use to discover model ids before passing one as the `model` argument to a predict tool. The same catalog is also available… List available models for a task. Use to discover model ids before passing one as the `model` argument to a predict tool. The same catalog is also available as the resource `gi://models`. Returns a FLAT object — {task, default_model, models: [...]} — not the {data, meta} envelope the predict tools return. Each model carries a `bio_spec`, whose useful fields are `request_max_bp` (the enforced ceiling, 500,000 everywhere) and `context_window_bp` (what the model reads in one step — compare your sequence length against it: a shorter one is scored against a padded window). `trained_window_bp` is the fixed receptive field where there is no sliding window (9,198 for g0-expression). `request_max_bp` is the only one of the three that is a cap; the window fields describe what the model scores, not what the route accepts.
load_demo_sequence
name*
Load a bundled demo reference sequence and return a handle. The server ships one curated, task-correct positive control per task (list them via the gi://sequences… Load a bundled demo reference sequence and return a handle. The server ships one curated, task-correct positive control per task (list them via the gi://sequences resource) — e.g. `expression_hbb_k562` is a ready-to-use K562 expression window for predict_expression. Stores the demo and returns a handle to pass to a predict_* tool: no Ensembl fetch, no quota. Handy for smoke-testing a prediction end-to-end.
predict_chromatin
read-only
modelsequencesequence_refsequence_name
Chromatin annotation across 919 features (G0 DeepSEA). 200–500,000 bp. The model reads a 1,000 bp context window; 200–999 bp is accepted and scored against a padde… Chromatin annotation across 919 features (G0 DeepSEA). 200–500,000 bp. The model reads a 1,000 bp context window; 200–999 bp is accepted and scored against a padded window.
predict_enhancer
read-only
modelsequencesequence_refsequence_name
Predict enhancer activity (G0 DeepSTARR). 50–500,000 bp. 50 bp is the task's admission floor (the API 422s below it), not a statement about what the model reads: e… Predict enhancer activity (G0 DeepSTARR). 50–500,000 bp. 50 bp is the task's admission floor (the API 422s below it), not a statement about what the model reads: enhancer models score a 249 bp context window, so 50–248 bp is accepted and scored against a padded window. For a meaningful call, submit at least the 249 bp context.
predict_expression
read-only
modelsequencetss_indexdescriptionsequence_refsequence_name
Predict a gene's expression from a TSS-centred window. Expression is cell-type-specific, so `description` (cell type / assay context, e.g. 'K562 cell line') is REQ… Predict a gene's expression from a TSS-centred window. Expression is cell-type-specific, so `description` (cell type / assay context, e.g. 'K562 cell line') is REQUIRED — the API rejects requests without it. The model scores exactly 9,198 bp centred on the TSS (±4,599). Two ways to supply that: - A sequence of exactly 9,198 bp already centred on the TSS. No `tss_index` needed — the midpoint is the only legal TSS. - A longer locus, 9,198–500,000 bp, plus `tss_index`: the 0-based offset of the TSS into it. The API cuts the window for you (sequence[tss_index-4599 : tss_index+4599]) and never scans for a TSS itself. Anything under 9,198 bp is rejected, here and by the API (422) — there is no padding or truncation fallback. `tss_index` is required for every other length, because a locus with no offset is indistinguishable from a mis-centred window. An offset that is merely WRONG (e.g. counted over a wrapped FASTA's characters, or against a chromosome coordinate instead of an offset into THIS sequence) still succeeds and scores the wrong window — verify meta.task_specific_counts.scored_window in the response. Easiest paths: fetch_gene_for_expression(gene) returns a ready-centred handle, and find_genes_and_predict_expression takes a raw region and finds each TSS for you.
predict_promoter
read-only
modelsequencesequence_refsequence_name
Predict promoter regions (G0). 300–500,000 bp. Returns the {data, meta} envelope: data.regions lists predicted promoters with start/end/score. 300 bp is t… Predict promoter regions (G0). 300–500,000 bp. Returns the {data, meta} envelope: data.regions lists predicted promoters with start/end/score. 300 bp is the task floor for every promoter model. The default g0-promoter-2000bp scans a 2,000 bp context window, so a shorter (but ≥300 bp) sequence is still scored — against a window padded out to that size. Check the chosen model's bio_spec.context_window_bp via list_models to know whether it saw real sequence or padding.
predict_splice
read-only
modelsequencesequence_refsequence_name
Predict splice donor/acceptor sites (G0 BigBird). 100–500,000 bp. The model reads a 15,000 bp context window, so anything shorter is scored against a padded window… Predict splice donor/acceptor sites (G0 BigBird). 100–500,000 bp. The model reads a 15,000 bp context window, so anything shorter is scored against a padded window — feed a whole transcript locus when you can. It is also strand-specific, and the wrong strand fails silently and plausibly — it returns sites at different positions, often still scoring above 0.9, not the near-zero scores once documented here. Nothing in the response flags it, so submit the transcript's own orientation (fetch_region takes `strand`).
store_inline_sequence
namesequence*