Build Databricks pipelines and Genie spaces on governed operational data
Genie answers in whatever definitions the lakehouse tables carry. When those tables are loaded from operational databases nobody documented, the answers inherit the guesswork. Datapace is built so the confirmed graph of your operational databases (entities, measures, grain, keys, PII flags, in your experts’ terms) is the semantic layer the pipeline and the Genie space consume: the pipeline mapping from operational tables into lakehouse tables is drafted from the graph, the table descriptions and metric definitions Genie reads are drafted from the same graph, and every draft is reviewed before it ships.
Today. Lakehouse tables loaded from operational databases by hand, described after the fact, with Genie answering in definitions nobody validated.
With Datapace. Pipeline mappings, table descriptions, and metric definitions drafted from the confirmed graph and reviewed by your team, so Genie answers in the terms your experts confirmed.
What Datapace does here
- The confirmed graph as the semantic layer
- Entities, measures, grain, keys, and relationships confirmed by your experts on the operational side are what the lakehouse side is built from, not rediscovered.
- Pipeline mapping drafted from the graph
- Datapace is built to draft the mapping from operational tables into lakehouse tables from the graph, including the joins the keys imply and the grain each table should keep.
- Descriptions and metrics Genie can use
- Table and column descriptions and metric definitions are drafted in your experts’ terms for the catalog Genie reads, so a question about revenue gets the confirmed definition of revenue.
- PII flags travel with the columns
- Columns your experts confirmed as personal data carry that flag into the pipeline draft, so masking and access are decided before the table exists.
- Reviewed before it ships
- Every draft, mapping or definition, is reviewed by your data team; the graph stays the source, so a change on the operational side surfaces as a proposed change downstream.
Questions teams ask
- Does Datapace replace the pipeline tooling?
- No. It drafts what the pipeline should do, from a graph your experts confirmed, and the pipeline runs on the tooling you already use in Databricks. What changes is where the definitions come from.
- Why does Genie need this?
- Genie is only as good as the descriptions and metric definitions on the tables it reads. When those are drafted from a confirmed graph of the operational source, in your experts’ terms, the answers are grounded in what the business means, not in a column name.
- What about the operational databases Databricks hosts itself?
- The same method applies to any operational database on any engine: Datapace reads it, infers meaning, has your experts confirm, and serves the graph. Where the database runs is a decision taken with each partner, not a constraint of the product.
See this on your own data
Bring a use case. We will show you what Datapace reads on your live database, what your experts would confirm, and what the agents would propose first.
Book a call