Semantic Layer for AI: Why Your AI Needs Business Definitions, Not Just Database Tables
What a semantic layer is, how metrics and dimensions work, and how it compares with databases, RAG and knowledge graphs when AI answers business questions.
Quick answer
A semantic layer is a governed translation layer between business language and physical data. It defines metrics (what is measured, such as net revenue), dimensions (how results can be sliced, such as region or product line), the joins and filters behind them, and who may see them. Applications ask for a metric by name; the layer generates the query.
For AI, the semantic layer replaces guesswork. Instead of asking a model to write SQL against hundreds of tables and hope it applies your exclusions correctly, the model selects from a short list of approved metrics and dimensions, and the layer computes the number the same way finance does.
Why database tables are not enough for AI
Databases store facts in a shape that suits the applications that write them. Answering a business question usually requires knowledge that is not in the schema: which status codes count as a completed order, which table holds the reporting currency, whether refunds reduce revenue in the month of sale or the month of refund, and which of three customer tables is authoritative.
A language model given the schema will produce plausible SQL. Plausible is the problem: a query can run, return a number and still use the wrong definition. Benchmarks such as BIRD were designed around this gap between understanding a question and querying a real, messy database; on that benchmark, human experts still score clearly higher than the best automated systems (BIRD benchmark). The fix is not only a better model; it is giving the model better-defined things to ask for. See text-to-SQL for business data for the query side.
Metrics, dimensions and semantic models
Most semantic layers share the same building blocks, even if each product names them differently.
| Building block | What it is | Example |
|---|---|---|
| Semantic model | A business view over one or more tables, with keys and joins | Orders model joining orders, lines and customers |
| Measure | An aggregation on a column | SUM(line_amount_net) |
| Metric | A named business calculation built from measures | Net revenue = gross sales minus discounts and refunds |
| Dimension | An attribute to group or filter by | Region, product category, sales channel |
| Time dimension | A date with grains and calendars | Order date by fiscal month |
| Entity / join key | How models connect | customer_id links orders to accounts |
| Access policy | Who can query what | Region managers see only their region |
metric: net_revenue
description: >
Revenue after discounts and refunds, excluding tax and
shipping. Owned by Finance. Used in board reporting.
type: derived
expr: gross_sales - discounts - refunds
filters:
- order_status in ('completed', 'partially_refunded')
- is_test_order = false
time_dimension: order_date # fiscal calendar, starts April
dimensions: [region, channel, product_category]
owner: finance-analytics
synonyms: ["revenue", "net sales"]Business definitions are the point
The value of a semantic layer is not the YAML. It is that a definition is agreed once, owned by someone and reused everywhere. Write a plain-language description for every metric and dimension, list synonyms people actually use, and document exclusions. Those descriptions are exactly what an AI model reads to choose the right metric, so they double as AI instructions.
This is one part of a wider business context layer, which also holds rules, entities, organizational structure, time and permissions that are not metrics.
How an AI assistant uses a semantic layer
A typical flow: the user asks a question; the assistant identifies candidate metrics and dimensions from their names, descriptions and synonyms; it fills a structured request (metric, dimensions, filters, time range); the semantic layer validates it, applies the user's access policy and generates SQL; the warehouse returns results; the assistant explains them and states which definitions it used. The model never writes free-form SQL against raw tables, which removes a whole class of errors and makes answers auditable.
User question ─▶ AI assistant
│ picks metric + dimensions
│ (structured request, not SQL)
▼
┌──────────────┐ definitions, joins,
│ Semantic │◀─ synonyms, owners,
│ layer │ access policies
└──────┬───────┘
│ generated SQL (governed)
▼
Warehouse / lakehouse
│ result rows
▼
Answer + "metric used: net_revenue (Finance)"Database vs semantic layer
The database stores and computes; the semantic layer decides what should be computed and how. You still need a well-modelled warehouse. The semantic layer sits on top and turns that model into business vocabulary.
| Database or warehouse | Semantic layer | |
|---|---|---|
| Unit | Tables, columns, rows | Metrics, dimensions, entities |
| Language | SQL | Business terms mapped to SQL |
| Definitions | Implicit in queries and reports | Explicit, named, owned, versioned |
| Consistency | Each query may differ | Same metric, same logic everywhere |
| AI interface | Model writes SQL against schema | Model selects governed metrics |
Semantic layer vs RAG vs knowledge graph
These three are often confused because all of them 'give AI context'. They work at different layers and answer different kinds of questions. A semantic layer works at the metrics layer over structured data. RAG works at the document layer, retrieving text passages. A knowledge graph works at the relationship layer, storing entities and the links between them (see knowledge graph vs vector database).
| Semantic layer | RAG | Knowledge graph | |
|---|---|---|---|
| Best question | How much? How many? Trend? | What does the policy say? | How is X connected to Y? |
| Data | Structured, aggregated | Unstructured text | Entities and relationships |
| Output | Computed numbers | Relevant passages | Paths, neighbours, facts |
| Correctness comes from | Governed definitions | Retrieval quality and citations | Modelled relationships |
| Typical failure | Missing metric or dimension | Wrong or stale passage | Incomplete or outdated graph |
Pro tip
Route by question type. Numeric business questions go to the semantic layer, policy and how-to questions go to retrieval, and relationship questions go to the graph. One assistant can use all three as tools.
Semantic layer options
Semantic layers come in three broad forms. Standalone or transformation-linked layers such as the dbt Semantic Layer (built on MetricFlow) and Cube define metrics once and serve many tools. Warehouse-native semantics such as Snowflake semantic views and Databricks Unity Catalog metric views store definitions as governed objects inside the platform, and both vendors position them as the basis for their natural-language analytics features. BI-native models such as LookML and Power BI semantic models live inside a BI tool.
Portability is improving. The Open Semantic Interchange (OSI) initiative, announced by Snowflake with Salesforce, dbt Labs, BlackRock, RelationalAI and others, is working on a vendor-neutral specification so semantic definitions can move between tools (dbt Labs on the OSI specification). It is young; check current tool support before depending on it.
AI agent use cases
Semantic layers are not only for chat-with-your-data. Agents use them as a safe, read-only analytics tool. Examples: a finance agent preparing a monthly variance commentary queries governed metrics and explains movements; a sales agent checks an account's revenue trend before drafting a renewal note; an operations agent monitors a fulfilment-rate metric and opens a ticket when it drops below threshold; an ecommerce merchandising assistant compares category margin across channels. In each case the agent calls a tool such as query_metric(metric, dimensions, filters, time_range), which is far easier to secure and evaluate than free-form SQL. See AI agent tool design.
How to implement a semantic layer for AI
- Start with the 10 to 20 metrics behind your most common business questions
- Agree definitions with their owners before writing any code
- Write descriptions and synonyms for every metric and dimension
- Choose the layer that fits your warehouse and BI stack
- Expose it to AI through a structured tool, not raw SQL access
- Apply row- and column-level access policies inside the layer
- Make the assistant cite the metric and definition it used
- Build an evaluation set of real questions with verified answers
- Treat metric changes like code changes: review, version, test
Common mistakes
Teams often expose every column as a dimension, which recreates the raw-schema problem with nicer names. Keep the vocabulary small and curated. Another mistake is skipping descriptions; for AI, the description is the interface. Some teams build a semantic layer for dashboards but let the AI assistant query raw tables anyway, which guarantees two sets of numbers. Finally, a semantic layer cannot answer questions it has no metric for; the assistant should say so rather than improvise a query.
Building AI analytics on company data?
ZSpace Labs connects AI assistants and agents to governed data through well-designed tools and APIs. See AI automation.
Conclusion
A semantic layer gives AI the one thing a schema cannot: agreed business meaning. It turns 'write some SQL' into 'choose from these governed metrics', which makes answers consistent with the rest of the business and easy to check. Use it for numeric questions, pair it with retrieval for documents and a graph for relationships, and treat its definitions as the shared language between people, BI tools and AI.
Common questions.
A semantic layer is a governed set of business definitions, mainly metrics and dimensions, mapped onto physical data. Tools and AI applications ask for 'net revenue by region last quarter' and the layer generates the correct query, so every consumer gets the same number.