This example provides DNA sequence search, supporting data file upload, feature extraction using Spark MLlib, and subsequent retrieval based on Milvus vector engine.
Engine Features
- Underlying feature vector similarity search
- Millisecond-level search on billions of records on a single server
- Near real-time search with distributed deployment support
- Insert, delete, search and update data at any time
DNA Background
Deoxyribonucleic acid (DNA) is one of the four major biological macromolecules found in living cells. DNA carries the genetic information necessary for synthesizing RNA and proteins, and is an essential biomolecule for organism development and normal functioning.
A DNA sequence refers to the primary structure of a DNA molecule, represented using a string of letters (A, T, C, G) that carries genetic information, either real or hypothetical. DNA sequencing methods include optical sequencing and chip-based sequencing.
Feature Extraction
Model training and inference use Spark MLlib:
Vector Engine Index Strategy
Function Introduction
- Input Address: http://localhost:8090
- Upload Data File:
1). Click the upload button to upload files.
2). Click the feature extraction button. Wait for file parsing, model training, feature extraction, and feature storage into the vector engine. Progress can be seen via the console.
- DNA Sequence Search — Enter text, click search, and see the returned list sorted by similarity.
ATGCCCCAACTAAATACTACCGTATGGCCCACCATAATTACCCCCATACTCCTTACACTATTCCTCATCACCCAACTAAAAATATTAAACACAAACTACCACCTACCTCCCTCACCAAAGCCCATAAAAATAAAAAATTATAACAAACCCTGAGAACCAAAATGAACGAAAATCTGTTCGCTTCATTCATTGCCCCCACAATCCTAG