PyPaimon Rust
This project builds the Rust-powered core for PyPaimon while also providing DataFusion integration for querying Paimon tables.
Install via PyPI:
pip install pypaimon-rust
If you want to use the native Python DataFusion SessionContext, install datafusion as well.
Query Paimon Tables with DataFusion
The recommended way to query Paimon tables is through SQLContext, which supports
multi-catalog registration, DDL, DML, and all Paimon-specific SQL extensions:
from pypaimon_rust.datafusion import SQLContext
ctx = SQLContext()
ctx.register_catalog("paimon", {
"warehouse": "/path/to/warehouse",
})
# DDL
ctx.sql("CREATE SCHEMA paimon.my_db")
ctx.sql("CREATE TABLE paimon.my_db.users (id INT, name STRING, PRIMARY KEY (id))")
# DML
ctx.sql("INSERT INTO paimon.my_db.users VALUES (1, 'alice'), (2, 'bob')")
# Query tables via SQL (catalog.database.table)
batches = ctx.sql("SELECT * FROM paimon.my_db.users")
Multimodal SQL Helpers
SQLContext also registers Python scalar UDFs for common BLOB media and vector
workflows. Install the optional media dependencies when you need image or video
decoding:
pip install "pypaimon-rust[video]"
batches = ctx.sql("""
SELECT
id,
media_info(content) AS info_json,
media_thumbnail(content, 160, 90) AS preview_png,
video_snapshot(content, 1000) AS frame_png
FROM paimon.my_db.assets
""")
For vector search, vector_from_json converts JSON-encoded embeddings into
Arrow List<Float32> values that can be used from a temporary table or subquery:
batches = ctx.sql("""
WITH queries AS (
SELECT id, vector_from_json(embedding_json) AS embedding
FROM paimon.my_db.query_embeddings
)
SELECT q.id AS query_id, r.id AS result_id
FROM queries q
CROSS JOIN LATERAL vector_search(
'paimon.my_db.items',
'embedding',
q.embedding,
10
) AS r
""")
Temporary Tables
You can register temporary in-memory tables programmatically. Names support the same resolution rules as SQL: bare names use the current catalog and database, partially qualified names use the current catalog, and fully qualified names specify catalog.database.table.
Register a single PyArrow RecordBatch as a temporary table:
import pyarrow as pa
batch = pa.record_batch([[1, 2], ["alice", "bob"]], names=["id", "name"])
ctx.register_batch("paimon.default.my_temp", batch)
batches = ctx.sql("SELECT * FROM paimon.default.my_temp")
# Drop it via SQL when no longer needed
ctx.sql("DROP TEMPORARY TABLE paimon.default.my_temp")
Alternatively, if you want to use the native Python DataFusion SessionContext,
install datafusion and register a PaimonCatalog:
from datafusion import SessionContext
from pypaimon_rust.datafusion import PaimonCatalog
catalog = PaimonCatalog({
"warehouse": "/path/to/warehouse",
})
ctx = SessionContext()
ctx.register_catalog_provider("paimon", catalog)
# Query tables via SQL (catalog.database.table)
df = ctx.sql("SELECT * FROM paimon.default.my_table LIMIT 10")
df.show()
REST Catalog
from datafusion import SessionContext
from pypaimon_rust.datafusion import PaimonCatalog
catalog = PaimonCatalog({
"metastore": "rest",
"uri": "http://localhost:8080",
"warehouse": "my_warehouse",
})
ctx = SessionContext()
ctx.register_catalog_provider("paimon", catalog)
Time travel queries are not supported in the Python binding at this time.