Skip to content

ZCaps for genomics

Compress massive genomic datasets without giving up selective access.

ZCaps brings adaptive lossless compression and chunk-addressable retrieval to large scientific datasets such as VCF, SAM, FASTQ, FASTA, GTF, GFF3 and related bioinformatics workloads.

SCIENTIFIC ARCHIVE98.90 GiB example
Integration maps the selected region to a byte range ↓
Selective chunk decodingOnly the two addressed chunks are decoded; the other eight remain compressed. Conceptual illustration.
Targeted 4 MiB extraction

Target the region. Decode the required chunks. Keep the rest compressed.

Lossless compression for genomic workflows

Compress

High-density lossless compression for highly structured scientific datasets.

Access

Retrieve targeted byte ranges by decoding only the relevant compressed chunks.

Scale

Independent chunks and files can be processed concurrently across server resources.

ZCaps addresses bytes and ranges. Chromosome or genomic-coordinate lookup requires an integration layer that maps those coordinates to byte ranges.

01 / Separate genomics benchmark campaign

Large datasets. Measured results.

Selected ZCaps results from the historical genomics campaign. These measurements use a different system and build from the current ZMDC benchmark.

VCF chr21

1000 Genomes phase 3

11.23 GiBoriginal dataset

134.6:1 at level 12

Compare levels 1, 6 and 12

Level 1

Compression ratio
119.5:1
Compression
1,958 MiB/s
4 MiB extraction
0.11 s

Level 6

Compression ratio
117.6:1
Compression
1,876 MiB/s
4 MiB extraction
0.11 s

Level 12

Compression ratio
134.6:1
Compression
1,488 MiB/s
4 MiB extraction
0.16–0.17 s

GTF

Gene annotation

4.45 GiBoriginal dataset

73.07:1 at level 12

Compare levels 1, 6 and 12

Level 1

Compression ratio
38.36:1
Compression
1,498 MiB/s
4 MiB extraction
0.15–0.16 s

Level 6

Compression ratio
55.76:1
Compression
1,544 MiB/s
4 MiB extraction
0.16–0.18 s

Level 12

Compression ratio
73.07:1
Compression
997 MiB/s
4 MiB extraction
0.22–0.25 s

SAM

Sequence alignment

30.93 GiBoriginal dataset

7.92:1 at level 12

Compare levels 1, 6 and 12

Level 1

Compression ratio
5.79:1
Compression
959 MiB/s
4 MiB extraction
0.27–0.29 s

Level 6

Compression ratio
6.79:1
Compression
808 MiB/s
4 MiB extraction
0.28–0.33 s

Level 12

Compression ratio
7.92:1
Compression
594 MiB/s
4 MiB extraction
0.42–0.46 s

FASTQ

Sequencing reads

6.34 GiBoriginal dataset

4.48:1 at level 12

Compare levels 1, 6 and 12

Level 1

Compression ratio
3.87:1
Compression
807 MiB/s
4 MiB extraction
0.35–0.37 s

Level 6

Compression ratio
4.20:1
Compression
697 MiB/s
4 MiB extraction
0.39–0.42 s

Level 12

Compression ratio
4.48:1
Compression
493 MiB/s
4 MiB extraction
0.63–0.67 s

GOA UniProt GAF

Functional annotation

98.90 GiBoriginal dataset

18.63:1 at level 12

Compare levels 1, 6 and 12

Level 1

Compression ratio
12.92:1
Compression
1,164 MiB/s
4 MiB extraction
0.302–0.315 s

Level 6

Compression ratio
16.25:1
Compression
1,041 MiB/s
4 MiB extraction
0.301–0.338 s

Level 12

Compression ratio
18.63:1
Compression
715 MiB/s
4 MiB extraction
0.437–0.488 s

0.11–0.75 smeasured targeted extraction envelope

30 formats · 3 compression levels · 90 compression runs · 270 targeted 4 MiB extraction runs · datasets up to 98.90 GiB.

SHA-256 validation of extracted ranges.

Benchmark system

Intel Core i9-13980HX
16 GiB RAM · NVMe SSD
Windows 11

Historical genomics campaign. Separate from the current ZMDC benchmark. Results are specific to the tested data, build and hardware.

Explore ZCaps technical evidence

02 / Scientific data compression

Different structures. Adaptive compression.

Maximum measured compression ratio · ZCaps level 12

ZCaps across different datasets. This chart is not a comparison against competing compressors.

03 / Selective access at scale

Access the part you need.

Observed targeted extraction remained sub-second across the tested size range.

GFF31.65 GiB
JSON6.09 GiB
VCF11.23 GiB
VCF23.50 GiB
Pfam38.71 GiB
FASTA64.20 GiB
GOA GPA71.88 GiB
GOA GAF98.90 GiB

0.11–0.75 sObserved campaign envelope · targeted 4 MiB extraction

File sizes shown are documented examples. The extraction envelope describes the campaign, not individual latency values for each file or a universal latency guarantee.

04 / Bioinformatics infrastructure

One layer for scientific data workflows.

Genomic archives

Retain byte-exact scientific records with lossless bioinformatics compression and selective retrieval.

Variant datasets

Explore VCF compression for large variant collections while retaining targeted byte-range access.

Sequencing pipelines

Integrate FASTQ compression and SAM compression into workflows for reads and alignments.

Research repositories

Reduce repository storage while preserving original files for reproducible research.

Large annotation datasets

Compress GTF, GFF3 and functional annotation records using a data-adaptive engine.

Local scientific infrastructure

Keep compression and genomic random access inside infrastructure controlled by your organization.

High-volume bioinformatics storage

Use scientific data compression across large collections, processing independent files concurrently.

Discuss your workload

Store more genomic data.
Access the part you need.