PBI-Scope Story — one read walkthrough¶
This page explains the full PBI-Scope flow in execution order.
For setup and commands, use the Installation Guide.
1) What PhageScope brings¶
PhageScope aggregates several public phage resources into consistent exports. PBI-Scope uses those exports as its main public phage input (metadata and FASTA).
2) Why Docker is central¶
PBI-Scope relies on Docker to keep paths, environments, and large intermediate data consistent.
pipelinebuilds dataanalysisreads data for notebooks/scriptsapiprovides REST API for exploration
Named volumes and bind mounts keep outputs persistent and auditable.
3) What the pipeline does (in order)¶
- Download public phage data from PhageScope sources
- Merge and normalize metadata with schema contracts
- Build merged FASTA files for phages and proteins + create indexes
- Validate private sources (if folders exist in
private_data/) - Prepare private mappings (private phages and hosts)
- Parse host fields from phage metadata
- Resolve hosts to NCBI assemblies
- Download host FASTAs from NCBI RefSeq
- Create DuckDB database and optimize analytical access
- Store reports and logs (validation, quality, failure logs)
- Build BLAST databases (phages, proteins, hosts, private, combined)
Expect long runtime on first execution
Steps 1–3, 7, 8, and 11 (downloading, file merging, host resolution, host download, and BLAST database building) are time-consuming on first run. The pipeline may appear stalled — especially during BLAST database building, which is just long — but it is performing I/O-heavy operations on large genomic datasets. These tasks typically take 10+ hours on first execution.
4) Resulting data product¶
After completion, PBI-Scope provides:
- DuckDB database for metadata exploration
- Indexed phage/protein FASTA files
- Host FASTA mapping for host retrieval
- Private phage mapping when private sources are present
- GFF3 gene annotations with index
- BLAST databases for sequence similarity searches (phages, proteins, hosts, private, combined)
- Pipeline logs and reports for traceability
5) How users work with it¶
The recommended interface is the analysis container with the pbi package.
- Use VS Code Dev Containers for full IDE workflow (preferred)
- Use Jupyter Lab for notebook-first workflow
- Use the REST API for quick exploration and metadata lookups
6) Where to go next¶
- Installation
- Analysis container
- Build Custom Containers
- API Reference
- Remote Access (API)
- Private data ingestion
- Notebooks README
7) Build your own environment¶
The default analysis container comes with Python and Jupyter Lab, but you are not limited to it. If you need a different setup — for example, R and Python together, or a container that runs scripts instead of notebooks — you can build your own.
A custom container connects to the same data volume as the default one, so it has full access to the database, sequences, and all other outputs. You write a simple Dockerfile, and Docker handles the rest. This is useful when you want to:
- Work in R alongside Python (e.g., for ggplot2 visualization or Bioconductor packages)
- Run automated scripts that execute without human interaction and save results to disk
- Use a different language like Julia, Rust, or anything else that can read DuckDB or FASTA files
- Share the database across multiple containers via the REST API, without loading data locally
The Building Custom Containers guide walks you through this step by step, with a ready-to-use R + Python example in the mount_scripts/ directory.