Stop building semantic layers in YAML
Data work is context-intensive. An agent needs to understand the shape of the data, metric definitions, join relationships, internal jargon, data quality issues, general business context, and more before it’s able to answer a data question correctly. And because context windows are finite, the information density of your context is a primary architectural concern.
Semantic layers are where the majority of this context lives. And right now, everyone building a semantic layer is choosing YAML: dbt’s Semantic Layer, Cube, Apache Ossie, and nearly all BI tools. YAML was fine back when end users didn’t read the semantic code directly, but in the world of agents, YAML’s verbosity is problematic.
In the agentic era, the language design constraints have changed. Not only does the language need to hold the semantic information; it also has to be a good API in itself for agents. One criteria is that it needs to be token efficient.
YAML vs. Graphene SQL
We designed our semantic layer language, Graphene SQL, with token efficiency in mind. Let’s look at some YAML-based semantic layers, and then compare them to Graphene SQL.
Here’s a small model with six columns, one derived dimension, one join, and four measures, defined in Cube:
cubes:
- name: orders
sql_table: orders
joins:
- name: users
sql: "{CUBE}.user_id = {users.id}"
relationship: many_to_one
dimensions:
- name: id
sql: id
type: number
primary_key: true
- name: created_at
sql: created_at
type: time
- name: status
sql: status
type: string
description: "One of 'Processing', 'Shipped', 'Complete', 'Cancelled', 'Returned'"
- name: amount
sql: amount
type: number
format: currency
currency: USD
description: "Amount paid by customer"
- name: cost
sql: cost
type: number
format: currency
currency: USD
description: "Cost of materials"
- name: revenue_recognized
sql: "{CUBE}.status IN ('Processing', 'Shipped', 'Complete')"
type: boolean
measures:
- name: revenue
sql: "CASE WHEN {revenue_recognized} THEN {CUBE}.amount ELSE 0 END"
type: sum
format: currency
currency: USD
description: "Amount paid by customer"
- name: cogs
sql: "CASE WHEN {revenue_recognized} THEN {CUBE}.cost ELSE 0 END"
type: sum
format: currency
currency: USD
description: "Cost of materials"
- name: profit
sql: "{revenue} - {cogs}"
type: number
format: currency
currency: USD
- name: profit_margin
sql: "(1.0 * {profit}) / {revenue}"
type: number
format: percent
And here’s the same model, defined in dbt’s Semantic Layer:
# models/staging/_orders.yml
semantic_models:
- name: orders
model: ref('stg_orders')
defaults:
agg_time_dimension: created_at
entities:
- name: order_id
type: primary
expr: id
- name: user
type: foreign
expr: user_id
dimensions:
- name: created_at
type: time
type_params:
time_granularity: day
- name: status
type: categorical
description: "One of 'Processing', 'Shipped', 'Complete', 'Cancelled', 'Returned'"
- name: revenue_recognized
type: categorical
expr: "status in ('Processing', 'Shipped', 'Complete')"
measures:
- name: revenue_raw
agg: sum
expr: "case when revenue_recognized then amount else 0 end"
agg_time_dimension: created_at
description: "Amount paid by customer"
- name: cogs_raw
agg: sum
expr: "case when revenue_recognized then cost else 0 end"
agg_time_dimension: created_at
description: "Cost of materials"
# models/metrics/_orders_metrics.yml
metrics:
- name: revenue
label: Revenue
type: simple
type_params:
measure: revenue_raw
meta:
currency: USD
- name: cogs
label: COGS
type: simple
type_params:
measure: cogs_raw
meta:
currency: USD
- name: profit
label: Profit
type: derived
type_params:
expr: "revenue - cogs"
metrics:
- name: revenue
- name: cogs
meta:
currency: USD
- name: profit_margin
label: Profit Margin
type: derived
type_params:
expr: "profit / revenue"
metrics:
- name: profit
- name: revenue
meta:
is_ratio: true
Now let’s see what this looks like in Graphene SQL.
table orders (
id BIGINT
user_id BIGINT
created_at DATETIME
status STRING -- One of 'Processing', 'Shipped', 'Complete', 'Cancelled', 'Returned'
amount FLOAT -- Amount paid by customer #currency=USD
cost FLOAT -- Cost of materials #currency=USD
join one users on user_id = users.id
revenue_recognized: status in ('Processing', 'Shipped', 'Complete')
revenue: sum(case when revenue_recognized then amount else 0 end) #currency=USD
cogs: sum(case when revenue_recognized then cost else 0 end) #currency=USD
profit: revenue - cogs #currency=USD
profit_margin: profit / revenue #ratio
)
That’s it! Here are the totals:
| lines | characters | ~tokens | |
|---|---|---|---|
| Cube (YAML) | 67 | 1,572 | ~395 |
| dbt Semantic Layer (YAML) | 75 | 1,747 | ~440 |
| Graphene SQL | 17 | 610 | ~150 |
This means that Graphene SQL is 2.6x more information-dense than Cube and 2.9x more than dbt. And it’s not so dense as to be illegible; in fact, it’s far easier to read, which is important for humans and agents alike.
What makes YAML SLs verbose
It isn’t any one vendor’s implementation choices. It comes down to two structural properties of YAML itself:
1. Repeated keys
Every field in a semantic layer YAML file repeats the same keys over and over (name:, type:, description:, etc.). A Graphene SQL field is one line: field_name: expression -- description #annotations. There’s no key to repeat, because position in the file is the schema, the same way amount FLOAT doesn’t need a type: FLOAT key beside it in a CREATE TABLE statement.
YAML’s object-list structure is built for arbitrary nesting; a semantic model doesn’t need arbitrary nesting. It needs a flat list of typed, named expressions, and SQL already has a compact, well-understood way to say that.
2. YAML doesn’t understand the SQL it’s holding
To a YAML parser, expr: "sum(case when revenue_recognized then amount else 0 end)" is an opaque string. It has no way to know that expression is an aggregation, so agg: sum has to be stated separately, even though the SQL already says so.
Graphene SQL doesn’t have this problem, because it owns its own SQL dialect end-to-end. The same parser that runs the query also reads the expression when the model is defined. It knows sum(...) makes a field a measure with no separate agg key needed, and for a handful of shapes it goes further: extract(year from dt) attaches timeGrain=year metadata automatically, with no annotation at all.
This matters now
Before AI, semantic layers were written by small data teams and never read directly by analyst users. A file’s real consumer was a BI tool’s frontend, which didn’t care about verbosity.
Agents care though. A 2-3x less compact semantic layer is 2-3x less room for what actually matters: the conversation; the dashboards it’s editing; the other tables it should have seen.
We think it’s time to move on from YAML. We hope you agree.
Want to see how good a state-of-the-art model is when armed with Graphene SQL? Drop us a line here.
