Compress
High-density lossless compression for highly structured scientific datasets.
ZCaps for genomics
ZCaps brings adaptive lossless compression and chunk-addressable retrieval to large scientific datasets such as VCF, SAM, FASTQ, FASTA, GTF, GFF3 and related bioinformatics workloads.
Target the region. Decode the required chunks. Keep the rest compressed.
High-density lossless compression for highly structured scientific datasets.
Retrieve targeted byte ranges by decoding only the relevant compressed chunks.
Independent chunks and files can be processed concurrently across server resources.
ZCaps addresses bytes and ranges. Chromosome or genomic-coordinate lookup requires an integration layer that maps those coordinates to byte ranges.
01 / Separate genomics benchmark campaign
Selected ZCaps results from the historical genomics campaign. These measurements use a different system and build from the current ZMDC benchmark.
1000 Genomes phase 3
134.6:1 at level 12
Gene annotation
73.07:1 at level 12
Sequence alignment
7.92:1 at level 12
Sequencing reads
4.48:1 at level 12
Functional annotation
18.63:1 at level 12
0.11–0.75 smeasured targeted extraction envelope
30 formats · 3 compression levels · 90 compression runs · 270 targeted 4 MiB extraction runs · datasets up to 98.90 GiB.
SHA-256 validation of extracted ranges.
Intel Core i9-13980HX
16 GiB RAM · NVMe SSD
Windows 11
Historical genomics campaign. Separate from the current ZMDC benchmark. Results are specific to the tested data, build and hardware.
Explore ZCaps technical evidence02 / Scientific data compression
Maximum measured compression ratio · ZCaps level 12
ZCaps across different datasets. This chart is not a comparison against competing compressors.
03 / Selective access at scale
Observed targeted extraction remained sub-second across the tested size range.
0.11–0.75 sObserved campaign envelope · targeted 4 MiB extraction
File sizes shown are documented examples. The extraction envelope describes the campaign, not individual latency values for each file or a universal latency guarantee.
04 / Bioinformatics infrastructure
Retain byte-exact scientific records with lossless bioinformatics compression and selective retrieval.
Explore VCF compression for large variant collections while retaining targeted byte-range access.
Integrate FASTQ compression and SAM compression into workflows for reads and alignments.
Reduce repository storage while preserving original files for reproducible research.
Compress GTF, GFF3 and functional annotation records using a data-adaptive engine.
Keep compression and genomic random access inside infrastructure controlled by your organization.
Use scientific data compression across large collections, processing independent files concurrently.
Discuss your workload