Vector Generation and Storage with BGE M3 and openGauss DataVec
BGE M3 is a multilingual, high-performance text embedding model developed by BAAI that converts text into semantically rich high-dimensional vector representations. This document focuses on BGE M3 and the vector database openGauss DataVec, and describes how to implement text vector generation and efficient storage. By combining these two tools, you can build more intelligent data retrieval and processing systems.
Note: For containerized deployment of openGauss DataVec, see the link.
Case 1: FlagEmbedding + openGauss DataVec
Environment Preparation
- Install dependencies
FlagEmbedding is a toolkit focused on retrieval-augmented large language models, providing various text embedding models and reranking models. Before using the bge-m3 model, you need to install this package first. For a detailed tutorial, refer to the Hugging Face official website.
pip3 install -U FlagEmbedding
pip3 install psycopg2- Load the bge-m3 model
In actual use, the FlagEmbedding framework supports automatically loading models from the Hugging Face model library by specifying the model name. The following supplements the steps for downloading the offline model. If you choose to load the model automatically, you can skip this section directly.
git lfs install ; git clone https://www.modelscope.cn/BAAI/bge-m3.gitNote that you need to install the Git LFS tool before you can download the complete model data. The download address is available on the git-lfs official website.
Practice
In the following example, we use the bge-m3 embedding model from FlagEmbedding to generate vector data and store it in the openGauss DataVec vector database.
import psycopg2
from psycopg2 import sql
from typing import List
import numpy as np
from FlagEmbedding import BGEM3FlagModel
def embedding(text):
model = BGEM3FlagModel(model_name_or_path = "BAAI/bge-m3")
sentence_vector_dict = model.encode(
text,
return_dense = True, # Set to return dense embedding, enabled by default
return_sparse = False, # Set to return sparse embedding, disabled by default
return_colbert_vecs = False # Set to return multi-vector (ColBERT), disabled by default
)
return sentence_vector_dict.get("dense_vecs")
def create_connection(dbname:str, user:str, password:str, host:str, port:int):
conn = psycopg2.connect(
dbname = dbname,
user = user,
password = password,
host = host,
port = port
)
cursor = conn.cursor()
return conn, cursor
def create_table(conn, cursor, table_name:str, dim:int):
cursor.execute(
sql.SQL(
"CREATE TABLE IF NOT EXISTS public.{table_name} (id BIGINT PRIMARY KEY, embedding vector({dim}));"
).format(table_name = sql.Identifier(table_name), dim = sql.Literal(dim))
)
conn.commit()
def insert(conn, cursor, table_name:str, embeddings:List[List[float]], ids:List[int]):
data = list(zip(ids, embeddings))
cursor.executemany(
sql.SQL("INSERT INTO public.{table_name} (id, embedding) VALUES(%s, %s);")
.format(table_name = sql.Identifier(table_name)), data
)
conn.commit()
print("Data inserted successfully!")
if __name__ == '__main__':
text = "openGauss is an open-source database"
emb = embedding(text)
dimensions = len(emb)
print("text : {}, embedding dim : {}, embedding : {} ...".format(text, dimensions, emb[:10]))
conn, cursor = create_connection("testdb", "test_user", YourPassword, "localhost", 5432)
create_table(conn, cursor, "test_table1", dimensions)
insert(conn, cursor, "test_table1", [emb.tolist()], [0])The output is as follows:
text : openGauss is an open-source database, embedding dim : 768, enbedding : [-0.05427849 -0.02701874 -0.05441538 0.0294214 -0.01936925 -0.00815862 0.01310737 -0.0480913 0.01261776 0.2954952] ...
Data inserted successfully.For details about using openGauss DataVec, see Python SDK Integration with Vector Database
Case 2: ollama + openGauss DataVec
Environment Preparation
- Load the model
For ollama installation, see openGauss-RAG Practice
ollama pull bge-m3- Verification
ollama list
NAME ID SIZE MODIFIED
bge-m3:latest 790764642607 1.2GB 18 minutes agoPractice
In the following example, we use the bge-m3 embedding model in ollama to generate vector data and store it in the openGauss DataVec vector database.
import ollama
import psycopg2
from psycopg2 import sql
from typing import List
def embedding(text):
vector = ollama.embeddings(model="bge-m3", prompt=text)
return vector["embedding"]
def create_connection(dbname:str, user:str, password:str, host:str, port:int):
conn = psycopg2.connect(
dbname = dbname,
user = user,
password = password,
host = host,
port = port
)
cursor = conn.cursor()
return conn, cursor
def create_table(conn, cursor, table_name:str, dim:int):
cursor.execute(
sql.SQL(
"CREATE TABLE IF NOT EXISTS public.{table_name} (id BIGINT PRIMARY KEY, embedding vector({dim}));"
).format(table_name = sql.Identifier(table_name), dim = sql.Literal(dim))
)
conn.commit()
def insert(conn, cursor, table_name:str, embeddings:List[List[float]], ids:List[int]):
data = list(zip(ids, embeddings))
cursor.executemany(
sql.SQL("INSERT INTO public.{table_name} (id, embedding) VALUES(%s, %s);")
.format(table_name = sql.Identifier(table_name)), data
)
conn.commit()
print("Data inserted successfully!")
if __name__ == '__main__':
text = "openGauss is an open-source database"
emb = embedding(text)
dimensions = len(emb)
print("text : {}, embedding dim : {}, embedding : {} ...".format(text, dimensions, emb[:10]))
conn, cursor = create_connection("testdb", "test_user", YourPassword, "localhost", 5432)
create_table(conn, cursor, "test_table1", dimensions)
insert(conn, cursor, "test_table1", [emb], [0])The output is as follows:
text : openGauss is an open-source database, embedding dim : 768, enbedding : [-0.5359194278717041, 1.3424185514450073, -3.524909734725952, -1.0017194747924805, -0.1950572431087494, 0.28160029649734497, -0.473337858915329, 0.08056074380874634, -0.22012852132320404, -0.9982725977897644] ...
Data inserted successfully.For details about how to use openGauss DataVec, see Python SDK Integration with Vector Database