Skip to content
All work

EVO2 VARIANT ANALYSIS·DESCI·BNB CHAIN

Genomic inference belongs off-chain. Trust belongs on-chain.

Evo2 Variant Analysis runs Google’s Evo 2 model on genomic data through Modal serverless GPUs, then anchors genome ownership, research votes, and a data marketplace on BNB Chain. It won the DeSci track at BNB Hack Kerala.

Role

Team lead / full-stack

Language

TypeScript, Python, Solidity

Model

Evo 2 (Arc Institute)

Chain

BSC Testnet

Evo2 Variant Analysis

Architecture

How the system fits together

Next.js client → FastAPI → Modal serverless GPU → BNB Chain contracts

The problem

Genomics has a compute-trust gap

Genomic datasets are large, inference is expensive, and the output has to be trustworthy enough for researchers to build on. Putting the raw compute on-chain is absurd. Putting nothing on-chain makes provenance impossible.

The question was: where does decentralization actually help, and where does it just add cost?

Competitive gap

What existing platforms get wrong

Centralized bio platforms make inference easy but lock the data and the provenance inside their database. Pure DeSci projects try to put everything on-chain, which is slow, expensive, and unnecessary.

Evo2 Variant Analysis keeps the heavy compute off-chain and uses the chain only for ownership, access control, and funding votes.

ApproachWhere it fails
Centralized platformData and provenance are trapped in one company’s database.
On-chain everythingCompute is too slow and storage is too expensive for genomics.
Evo2 Variant AnalysisCompute off-chain, trust anchors on-chain.

The split

What goes on-chain and what does not

We run Evo 2 inference on Modal’s serverless GPUs because that is where the FLOPs are cheap and elastic. The model output is stored off-chain, but a hash of the output goes into the genome NFT metadata. A marketplace contract gates access based on ownership, and a DAO contract records votes on which studies get funded.

The chain never sees the genome. It sees ownership, consent, and money.

What goes on-chain and what does not
ConcernLayerWhy
GPU inferenceModal serverlessOn-chain compute cannot handle Evo 2.
Result storageOff-chain + hashPrivacy and size; hash keeps it reproducible.
OwnershipBNB Chain NFTProvenance without exposing the data.
Access controlMarketplace contractPay before download, enforced by chain.
Research fundingDAO contractTransparent, on-chain voting.

From sample to NFT

The inference pipeline

A user uploads a genomic sample. FastAPI validates the file and routes it to a Modal function that runs Evo 2. The returned variant, confidence, and raw hash are written to the genome NFT metadata. The raw output is available only to the owner and whoever they grant access to.

api/modal/analyze.py
@modal.function(gpu='A100')
def analyze_variant(sample_path: str) -> Analysis:
    result = evo2.run(sample_path)
    return Analysis(
        variant=result.variant,
        confidence=result.confidence,
        hash=sha256(result.raw).hexdigest(),
    )

Results

What the numbers say

During the hackathon demo we ran synthetic analyses through the pipeline and compared end-to-end latency against a naive always-on GPU setup. The modal cold start is acceptable because inference dominates the runtime.

−40%

Inference latency

VS ALWAYS-ON GPU

1,000+

Analyses run

DEMO PERIOD

100+

Genome NFTs minted

TESTNET

Source · hackathon demo, synthetic samples

Lessons

What building it taught me

Decentralization is a seasoning, not a main ingredient. The winning move was to use the chain only where it created trust — ownership, voting, payments — and keep the heavy science in the cloud.

Bio researchers do not care about the consensus mechanism. They care that the result is reproducible and that their data is not trapped in someone’s database.

The chain should record who owns the science, not try to do the science.

Still open

Where it is still rough

Testnet only

The contracts are deployed on BSC Testnet. Mainnet gas costs and storage economics would change the pricing model.

Synthetic data

Demo analyses used synthetic genomes with known variants. Real samples have noise, consent layers, and regulatory questions we did not model.

Access durability

If the off-chain storage provider disappears, the NFT still owns a hash of a file nobody can fetch. A real system needs pinning incentives or self-hosting.

In short

What Evo2 came down to

01

Compute off-chain

The model is too expensive and too private to run on-chain.

02

Anchor trust on-chain

Ownership, consent, and payment are the parts that benefit from being public and immutable.

03

Hash the result

A hash in the NFT makes the analysis reproducible without exposing the genome.

More work