From biological truth to deployable detection intelligence.
BIONET connects qualified biological references, reproducible detection analytics, and federated execution so the same validated intelligence can operate across independently governed environments.
FDA-ARGOS: What Are We Detecting?
BIONET begins with biological truth. FDA-ARGOS is the proposed reference and quality-control foundation against which new observations can be compared with regulatory-grade confidence.
Every detection system is constrained by its reference universe. A sequencing algorithm can produce an extremely precise answer to the wrong question if the reference genome is misidentified, contaminated, poorly assembled, incompletely annotated, or missing relevant diversity. Public sequence repositories are indispensable, but their contents vary in quality, provenance, metadata completeness, and suitability for regulatory or operational decisions. BIONET therefore requires a curated trust layer rather than relying on unqualified public data alone.
FDA-ARGOS, the FDA dAtabase for Reference Grade micrObial Sequences, was established as a public collection of quality-controlled and curated microbial genomic data to support research, diagnostic development, in-silico validation, and regulatory decision-making. Official FDA materials describe the database as a collaborative resource designed to help advance infectious-disease next-generation sequencing. Within BIONET, its role expands naturally from regulatory reference resource to the biological baseline for distributed surveillance.
Regulatory-grade does not mean merely “high quality” in a general sense. It implies a documented chain of confidence: independent organism identification, high-quality sequencing and assembly, adequate coverage and depth, relevant metadata, provenance linking the isolate or specimen to the sequence, and quality-control attributes that allow downstream users to determine whether a genome is fit for a defined purpose. These requirements reduce the risk that detection confidence is driven by reference artifacts rather than true biology.
| Element | Why it matters |
|---|---|
| Curated organism identity | Establishes confidence that a reference sequence represents the named organism or strain. |
| Assembly and coverage QC | Reduces false variation and structural signals caused by fragmented or erroneous assemblies. |
| Metadata and provenance | Preserves the context needed to interpret host, geography, collection, method, and lineage. |
| Functional annotation | Allows genomic observations to be mapped to virulence, resistance, regulatory, mobile, and structural elements. |
| Qualification tools | Creates a scalable path for mining public and partner data rather than resequencing every organism. |
The supplied program materials report that FDA-ARGOS currently includes more than 10,000 assemblies, including +3000 bacterial, +7000 viral, and ~100 fungal complete genomes. The same materials estimate that current coverage represents roughly 21 percent of recognized human pathogens and frame the primary limitation as quantity rather than credibility.
The proposed expansion strategy is not limited to generating new sequences. ARGOS-QC tools are intended to mine public or partner repositories and identify assemblies that satisfy FDA-ARGOS requirements. Official FDA descriptions similarly emphasize the need for quality matrices, scoring approaches, and public-database mining to expand the resource more sustainably. This is strategically important because a global monitoring system must grow faster than a laboratory-by-laboratory resequencing model can support.
BIONET also requires a library of functional pieces, not only whole genomes. The relevant objects include pathogenicity islands, antimicrobial-resistance genes, toxin and virulence determinants, mobile genetic elements, promoters and other regulatory regions, conserved protein domains, host-interaction elements, and synthetic-biology motifs. HIVE can map observed coverage and variation to these elements, but ARGOS and associated knowledge bases provide the trusted reference coordinates, annotations, and biological priors that make such interpretation credible.
BIONET intends expansion of FDA-ARGOS toward approximately 90 percent coverage of recognized human pathogens. That estimate includes pathogen-database expansion, regulatory pipelines, data infrastructure, and program management. Within the BIONET narrative, the investment logic is sequential: expand reference truth, operationalize detection against those references, and continuously improve both through field observations.
DNAHIVE Chief Scientist Dr. Vahan Simonyan leads the current FDA-ARGOS contract as Principal Investigator. FDA identifies Agreement No. 75F40121C00167 over five years beginning in April 2021, and describes the effort as enabling FDA and industry to use harmonized, well-characterized datasets for development and evaluation of pathogen-detection devices.
Dr. Simonyan’s role is strategically relevant because the BIONET concept depends on continuity across the reference, analytics, and federation layers. The supplied biography describes his prior service as a senior scientist at NIH/NCBI, a genetics lead and Director of R&D Bioinformatics at FDA, and the scientist responsible for the development, establishment and operations of FDA-HIVE since 2010. In BIONET, the same scientific leadership connects the creation of trusted microbial references to the design of the analytic environment that consumes them.
FDA-HIVE/H2O: How Are We Detecting It?
HIVE is the analytic engine that converts raw genomic and multi-omic observations into reproducible evidence about organism identity, variation, function, evolution, and anomaly.
Once a trusted reference layer exists, the next challenge is computational. A modern biological sample can contain millions or billions of sequencing reads from multiple organisms, hosts, contaminants, and background species. The important signal may be a low-abundance pathogen, a divergent strain, a structural rearrangement, a small subpopulation, an unmapped fragment, or an expressed protein that changes the risk interpretation. No single alignment or classifier is sufficient.
The High-performance Integrated Virtual Environment, or HIVE, was originally developed as a secure distributed storage and computing environment for next-generation sequencing. Official FDA materials describe it as a multicomponent infrastructure providing authorized users with secure web access to deposit, retrieve, annotate, compute on, and visualize large NGS datasets. BIONET exploits the current HIVE/H2O platform as a production-grade environment spanning genomics, multi-omics, imaging, clinical metadata, massively parallel analytics, provenance, security controls, APIs, and role-based visualization.
The regulatory context is a major differentiator. The FDA states that HIVE has maintained an FDA Authorization to Operate since 2012 and characterizes it as the ATO-ed platform at FDA fit for maintaining and computing on large-scale multi-omics datasets. They further report that FDA has used HIVE for research and regulatory review since 2011, with more than 27 petabytes of genomic information across 20 deployed U.S. instances.
FDA describes HIVE as supporting the genomic core and regulatory work across cell and gene therapies, vaccines, diagnostics, microbial therapeutics, oncology, hematology, CRISPR-based products, T-cell therapies, AAV and protein-delivery systems, pathogen detection, and vaccine safety. That breadth matters to BIONET because the platform is not a narrow pathogen classifier. It is a regulated computational ecosystem designed to maintain very large data assets, execute versioned pipelines, preserve provenance, and expose results through scientific interfaces.
BIONET’s HIVE detection stack is layered to answer the following questions.
- Which organisms are present? Identity.
- How different is this organism from what we knew?
- How does the pathogen impact us? Function.
- Has this pathogen been altered?
- Is the pathogen diversifying itself to avoid detection?
- What is the dark matter of completely unknown species?
- Are the pathogens active?
- Are our detection signals part of normal evolution or adversarial design?
| HIVE capability | BIONET function |
|---|---|
| High-performance storage and compute | Processes extra-large sequencing and multi-omic datasets with parallel execution. |
| Pipeline and application framework | Packages repeatable analytic methods, dependencies, references, and parameters. |
| Provenance and audit controls | Records inputs, methods, versions, intermediate results, and outputs for review. |
| Scientific visualization | Allows experts to inspect taxonomic, coverage, variation, structural, and population-level evidence. |
| Role-based interfaces | Translates raw analytics into technician, analyst, expert, and command-level outputs. |
| Federation APIs | Allows HIVE applications and data services to be deployed through the FEAST execution fabric. |
BIONET’s HIVE detection stack is layered to answer the following questions.
Which organisms are present?
The first layer asks which organisms are present. BIONET identifies HIVE-Censuscope, Pathoscope, Kraken 2, and MetaPhlAn4 as complementary approaches. Censuscope iteratively subsamples and maps reads until diversity estimates converge across taxonomic levels. Pathoscope uses Bayesian inference to account for sequence quality and mapping uncertainty. Kraken 2 provides rapid k-mer classification. MetaPhlAn4 uses clade-specific marker genes. The operational choice can vary by organism class, sequencing modality, speed, sensitivity, and required taxonomic resolution.
How different is this organism from what we knew?
Species detection is followed by high-sensitivity characterization. HIVE-Hexagon is described as a read aligner optimized for diverse viral and bacterial genomes, including more distant homology than conventional human-genome-oriented short-read tools may capture. HIVE-Heptagon then produces coverage and pileup information, point mutations, insertions, deletions, and structural variation. These layers allow BIONET to distinguish a simple organism hit from a biologically meaningful deviation.
How does the pathogen impact us?
Functional annotation converts sequence differences into biological interpretation. HIVE adapted use of NCBI, EBI, PIR, and UniProt resources to map coverage and variation to protein domains and functional elements. A missing pathogenicity island, an inserted conserved domain, a mobile resistance element, or an altered regulatory region can therefore be evaluated differently from a neutral sequence change. Species-specific biological priors and subject-matter expertise remain essential because the same genetic event may have different implications in different organisms.
Has this pathogen been altered?
Structural and recombinant detection is central to the adversarial mission. HIVE-Nonagon, also described as a Defective Viral Genome Profiler, traces split or chimeric alignments across different genomic regions or species. It is intended to identify translocations, copy-number changes, strand reversals, cross-species recombination, unexpected functional insertions, and sequence jumps into motifs associated with synthetic biology. These signals do not prove intent, but they create high-value indicators for expert review and attribution analysis.
Is the pathogen diversifying itself to avoid detection?
Clonal diversity provides another early-warning channel. RNA viruses and many bacterial pathogens exist as populations of related variants rather than a single genome. Conventional consensus assemblies can collapse that diversity and discard weak but important signals. HIVE-Hexahedron is described as assembling multiple related genomic trajectories into a “Nephosome,” preserving the structure and abundance of quasi-species. This supports detection of emerging variants, unusual diversification rates, selection pressure, and subpopulation patterns inconsistent with expected evolution.
What is the dark matter of completely unknown species?
Reference-based analysis is deliberately followed by de novo analysis of unmapped reads. Most unmapped content in environmental and clinical samples will be ordinary regional biological diversity. A smaller subset may form contigs or scaffolds with open reading. Frames and conserved domains related to pathogen families. HIVE-adapted de novo tools, combined with FDA-ARGOS quality-control protocols, can flag high-quality assemblies for resampling and expert assessment. Over time, region-specific background libraries can help separate normal “dark matter” from truly novel signals.
Are the pathogens active?
The metagenomics/metaproteomics concept adds functional confirmation. Sequencing can reveal that a toxin, resistance gene, promoter, or unusual splice is genetically present; proteomics can help determine whether the relevant protein is expressed and at what level. This distinction between latent capability and active biological threat is particularly important for bacteria, fungi, yeasts, protozoa, and other complex samples where genomic potential alone may overstate operational risk.
Are our detection signals part of normal evolution or adversarial design?
HIVE also provides the feature space for longitudinal AI. H2O supports three-dimensional operational representation of time, location, and genomic signal, supported by functional discriminant models, Fourier and wavelet methods, spatial-temporal diffusion models, continuous hidden Markov models, recurrent neural networks, and prospective genomic foundation models. The purpose is not to replace biological analysis with a single black box. It is to learn expected behavior, define statistically bounded baselines, and flag low-probability deviations for human and operational review.
| Analytic layer | Primary question answered | Illustrative output |
|---|---|---|
| Taxonomic census | What organisms are present? | Species/clade identities, abundance, diversity, co-occurrence. |
| Coverage and variation | How does the observed genome differ? | SNVs, indels, structural changes, coverage gaps, entropy. |
| Functional annotation | What might the differences do? | Virulence, AMR, toxin, regulatory, mobile, and host-interaction implications. |
| Recombinant detection | Are biological pieces arranged abnormally? | Chimeric reads, translocations, cross-species joins, synthetic motifs. |
| Clonal diversity | What is changing below the consensus? | Subpopulation trajectories, selection, emerging variants, diversification rate. |
| De novo discovery | What remains outside known references? | Contigs, ORFs, conserved domains, candidate novel organisms. |
| Proteomics and expressions | Is the species expressing pathogenicity? | Expression profiling, proteomics. |
| Longitudinal AI | Is behavior abnormal across time and place? | Anomaly scores, predicted ranges, low-probability deviations, propagation alerts. |
The final product is not a raw bioinformatics report. The BIONET portal concept separates technician triage, sample inventory, analyst diversity views, administrative controls, expert evidence review, and command reporting.
The signals on sample analysis can have different status: gray for processing, green for expected baseline, orange for elevated but explainable activity, red for unexpected pathogen or functional/propagation signal requiring attention, and blue for sample or data-quality failure.
This structure shortens the path from computation to action while preserving access to the underlying evidence.
HIVE IN ONE SENTENCE FDA-HIVE/H2O transforms sequencing and multi-omic data into a versioned, auditable chain of evidence about identity, variation, function, evolution, and anomaly, then presents that evidence at the level required by technicians, scientists, regulators, and commanders.
ARPA-H BDF FEAST: Where Are We Detecting?
FEAST, Federated Ecosystem for Analytics and Standardization Technologies, is the architecture that allows the same certified analytic logic to operate across heterogeneous data sources respecting security and governance, without requiring centralized data aggregation.
FEAST was developed under the ARPA-H Biomedical Data Fabric program. The official ARPA-H program describes a national need to connect biomedical data from thousands of sources, overcome incompatible data dialects, improve provenance and reproducibility, enable multi-source AI/ML, and maintain privacy and security. The FEAST architecture translates those goals into two core technical ideas: agnostic federation and agnostic harmonization.
Agnostic federation moves computation to data instead of moving data to computation. Agnostic harmonization allows computers to discover what standards, protocols, vocabularies, and permissions are available at a site, then apply the transformations required by a validated application at the point of use. Together, these ideas allow BIONET to work with existing systems rather than requiring every participant to adopt one database, one schema, one cloud provider, or one governance model.
The complete FEAST fabric includes a Constitution layer of protocols, a network of participating nodes, one or more physical or virtual FEAST units at each node, application and virtual-machine repositories, reference knowledge bases, a Library of Transformation modules, BioCompute certification, search and query services, controllers, a Compute Motility Orchestrator, resource virtualization, encryption and signaling objects, research dashboards, and a control center. Each layer has a distinct role in making distributed execution predictable.
OPERATIONAL RESPONSIBILITY INSIDE BIONET
| FEAST component | Operational responsibility in BIONET |
|---|---|
| FEAST Constitution | Defines high-level handshake, type, communication, permission, and interoperability protocols. |
| Node and FEAST unit | Creates a governed execution enclave inside or adjacent to each participating organization. |
| Controller | Coordinates search, application launch, resources, certificates, job status, and communication with other nodes. |
| Compute Motility Orchestrator | Starts, suspends, transfers, resumes, and terminates virtualized analytic processes. |
| Resource virtualization | Captures file, API, and database access and redirects the application to authorized local resources. |
| Library of Transformations | Performs semantic mapping, syntactic repackaging, ontology translation, normalization, and decompression at runtime. |
| VM and application repositories | Distribute preconfigured analytic environments with correct code, dependencies, references, and versions. |
| BioCompute certification | Describes and validates applications for transparency, reproducibility, risk assessment, and controlled execution. |
| SIGO and access controls | Protect sensitive values and permit authorized re-identification only within the originating site. |
The FEAST Constitution is particularly important. It is not a single data standard. It is a protocol for discovering and communicating which standards a source and an application support, which transformations are available, which permissions apply, and how the requested data can be delivered correctly. This allows harmonization at the point of use. An organization does not need to replace its internal databases, and a validated HIVE application does not need to be rewritten for every site.
The Library of Transformations operationalizes this approach. LOT modules can translate semantic concepts, syntactic packaging, vocabularies, ontologies, units, coding systems, genomic formats, compression schemes, API conventions, and database dialects. Examples in the supplied architecture include FASTQ/FASTA and SAM transformations, runtime decompression, SQL-flavor harmonization, FHIR packaging, and FHIR-to-OMOP conversion. Transformation chains can be assembled when no single module connects the source and target representation.
This is a crucial distinction from one-time data harmonization projects. Traditional integration requires every source to transform and maintain a duplicate standardized dataset before analysis can begin. FEAST encodes transformations once, validates them, and applies them as needed. New source standards and new application requirements can be added to the library without redesigning the entire network. For BIONET, this makes onboarding a new laboratory or surveillance domain a configuration and validation problem rather than a wholesale data-migration program.
Compute motility is the second defining capability. A user or automated surveillance process first performs a distributed search that returns metadata about available data, not the data themselves. A validated HIVE application is selected from the repository, together with required ARGOS references and dependencies. The FEAST controller verifies the BioCompute certificate, instantiates the virtual processes locally, maps authorized resources, and starts execution. When the application requires a resource available only at another node, the process can be suspended, transferred, and resumed inside the destination node’s enclave.
Only the application state and authorized supporting assets move. Site data, staging stores, local controller state, and protected source systems do not migrate. From the application’s perspective, the required resource becomes available through a virtualized file system, socket, or database interface. From the institution’s perspective, the computation enters a controlled enclave, accesses only pre-approved resources, and produces only defined outputs.
This operating model is referred to as Motile Intelligence. The intelligence is motile because the algorithmic capability, model state, reference assets, and validated process move across the network while sensitive biological data remain under local governance. A HIVE detection workflow can therefore analyze a hospital, a wastewater laboratory, an agricultural site, an overseas partner, and a defense laboratory using the same analytic definition without copying all source data into one repository.
Security is architectural rather than an afterthought. FEAST units are deployed as on-premises appliances or virtual units inside a private network. Firewalls restrict traffic to known FEAST peers and certified repositories over approved channels. Resource virtualization blocks unauthorized file, port, IP, API, and database access. Site-specific data inventories define which resources, protocols, authorization methods, and de-identification procedures are available. Temporary staging areas are wiped between analytic sessions.
FEAST architecture also implements SIGO (Signaling Objects), a runtime encryption and de-identification paradigm in which protected values can be reversibly re-identified only by authorized users and only in between encrypted memory and CPU, not data source, at the originating site. When an application reaches a protected resource, the source metadata can guide the computation to the node where decryption and permitted use are possible. Results do not even need to be re-encrypted before they are exposed because only the CPU accesses the unencrypted data, not movable memory.
Auditability supports both governance and incident response. BIONET and FEAST materials describe logging at application, system, database, and user levels, with timestamps, identity, actions, results, permissions, derived outputs, and configuration versions. BioCompute objects capture process rationale, datasets, code and dependency versions, and expected behavior. The goal is a reconstructable chain of evidence: what ran, where it ran, by whom, what it accessed, what transformations were applied, and what result was produced.
Only the application state and authorized supporting assets move. Site data, staging stores, local controller state, and protected source systems do not migrate. From the application’s perspective, the required resource becomes available through a virtualized file system, socket, or database interface. From the institution’s perspective, the computation enters a controlled enclave, accesses only pre-approved resources, and produces only defined outputs.
This operating model is referred to as Motile Intelligence. The intelligence is motile because the algorithmic capability, model state, reference assets, and validated process move across the network while sensitive biological data remain under local governance. A HIVE detection workflow can therefore analyze a hospital, a wastewater laboratory, an agricultural site, an overseas partner, and a defense laboratory using the same analytic definition without copying all source data into one repository.
Security is architectural rather than an afterthought. FEAST units are deployed as on-premises appliances or virtual units inside a private network. Firewalls restrict traffic to known FEAST peers and certified repositories over approved channels. Resource virtualization blocks unauthorized file, port, IP, API, and database access. Site-specific data inventories define which resources, protocols, authorization methods, and de-identification procedures are available. Temporary staging areas are wiped between analytic sessions.
FEAST architecture also implements SIGO (Signaling Objects), a runtime encryption and de-identification paradigm in which protected values can be reversibly re-identified only by authorized users and only in between encrypted memory and CPU, not data source, at the originating site. When an application reaches a protected resource, the source metadata can guide the computation to the node where decryption and permitted use are possible. Results do not even need to be re-encrypted before they are exposed (because only CPU accesses the unencrypted data, not movable memory).
Auditability supports both governance and incident response. BIONET and FEAST materials describe logging at application, system, database, and user levels, with timestamps, identity, actions, results, permissions, derived outputs, and configuration versions. BioCompute objects capture process rationale, datasets, code and dependency versions, and expected behavior. The goal is a reconstructable chain of evidence: what ran, where it ran, by whom, what it accessed, what transformations were applied, and what result was produced.
FEAST IN ONE SENTENCETFEAST enables the same validated HIVE/H2O analytics package to execute across independently governed sites by moving certified computation to data, standardizing at runtime, and returning authorized results with provenance.