Graphene

Stop building semantic layers in YAML

September 8, 2026·Kevin Marr

Ivan Žabota, In the Field (1895–1915), Slovak National Gallery — via Europeana on Unsplash

Data work is context-intensive. An agent needs to understand the shape of the data, metric definitions, join relationships, internal jargon, data quality issues, general business context, and more before it’s able to answer a data question correctly. And because context windows are finite, the information density of your context is a primary architectural concern.

Semantic layers are where the majority of this context lives. And right now, everyone building a semantic layer is choosing YAML: dbt’s Semantic Layer, Cube, Apache Ossie, and nearly all BI tools. YAML was fine back when end users didn’t read the semantic code directly, but in the world of agents, YAML’s verbosity is problematic.

In the agentic era, the language design constraints have changed. Not only does the language need to hold the semantic information; it also has to be a good API in itself for agents. One criteria is that it needs to be token efficient.

YAML vs. Graphene SQL

We designed our semantic layer language, Graphene SQL, with token efficiency in mind. Let’s look at some YAML-based semantic layers, and then compare them to Graphene SQL.

Here’s a small model with six columns, one derived dimension, one join, and four measures, defined in Cube:

cubes:
  - name: orders
    sql_table: orders

    joins:
      - name: users
        sql: "{CUBE}.user_id = {users.id}"
        relationship: many_to_one

    dimensions:
      - name: id
        sql: id
        type: number
        primary_key: true

      - name: created_at
        sql: created_at
        type: time

      - name: status
        sql: status
        type: string
        description: "One of 'Processing', 'Shipped', 'Complete', 'Cancelled', 'Returned'"

      - name: amount
        sql: amount
        type: number
        format: currency
        currency: USD
        description: "Amount paid by customer"

      - name: cost
        sql: cost
        type: number
        format: currency
        currency: USD
        description: "Cost of materials"

      - name: revenue_recognized
        sql: "{CUBE}.status IN ('Processing', 'Shipped', 'Complete')"
        type: boolean

    measures:
      - name: revenue
        sql: "CASE WHEN {revenue_recognized} THEN {CUBE}.amount ELSE 0 END"
        type: sum
        format: currency
        currency: USD
        description: "Amount paid by customer"

      - name: cogs
        sql: "CASE WHEN {revenue_recognized} THEN {CUBE}.cost ELSE 0 END"
        type: sum
        format: currency
        currency: USD
        description: "Cost of materials"

      - name: profit
        sql: "{revenue} - {cogs}"
        type: number
        format: currency
        currency: USD

      - name: profit_margin
        sql: "(1.0 * {profit}) / {revenue}"
        type: number
        format: percent

And here’s the same model, defined in dbt’s Semantic Layer:

# models/staging/_orders.yml
semantic_models:
  - name: orders
    model: ref('stg_orders')
    defaults:
      agg_time_dimension: created_at
    entities:
      - name: order_id
        type: primary
        expr: id
      - name: user
        type: foreign
        expr: user_id
    dimensions:
      - name: created_at
        type: time
        type_params:
          time_granularity: day
      - name: status
        type: categorical
        description: "One of 'Processing', 'Shipped', 'Complete', 'Cancelled', 'Returned'"
      - name: revenue_recognized
        type: categorical
        expr: "status in ('Processing', 'Shipped', 'Complete')"
    measures:
      - name: revenue_raw
        agg: sum
        expr: "case when revenue_recognized then amount else 0 end"
        agg_time_dimension: created_at
        description: "Amount paid by customer"
      - name: cogs_raw
        agg: sum
        expr: "case when revenue_recognized then cost else 0 end"
        agg_time_dimension: created_at
        description: "Cost of materials"

# models/metrics/_orders_metrics.yml
metrics:
  - name: revenue
    label: Revenue
    type: simple
    type_params:
      measure: revenue_raw
    meta:
      currency: USD

  - name: cogs
    label: COGS
    type: simple
    type_params:
      measure: cogs_raw
    meta:
      currency: USD

  - name: profit
    label: Profit
    type: derived
    type_params:
      expr: "revenue - cogs"
      metrics:
        - name: revenue
        - name: cogs
    meta:
      currency: USD

  - name: profit_margin
    label: Profit Margin
    type: derived
    type_params:
      expr: "profit / revenue"
      metrics:
        - name: profit
        - name: revenue
    meta:
      is_ratio: true

Now let’s see what this looks like in Graphene SQL.

table orders (
  id BIGINT
  user_id BIGINT
  created_at DATETIME
  status STRING -- One of 'Processing', 'Shipped', 'Complete', 'Cancelled', 'Returned'
  amount FLOAT -- Amount paid by customer #currency=USD
  cost FLOAT -- Cost of materials #currency=USD

  join one users on user_id = users.id

  revenue_recognized: status in ('Processing', 'Shipped', 'Complete')

  revenue: sum(case when revenue_recognized then amount else 0 end) #currency=USD
  cogs: sum(case when revenue_recognized then cost else 0 end) #currency=USD
  profit: revenue - cogs #currency=USD
  profit_margin: profit / revenue #ratio
)

That’s it! Here are the totals:

lines characters ~tokens
Cube (YAML) 67 1,572 ~395
dbt Semantic Layer (YAML) 75 1,747 ~440
Graphene SQL 17 610 ~150

This means that Graphene SQL is 2.6x more information-dense than Cube and 2.9x more than dbt. And it’s not so dense as to be illegible; in fact, it’s far easier to read, which is important for humans and agents alike.

What makes YAML SLs verbose

It isn’t any one vendor’s implementation choices. It comes down to two structural properties of YAML itself:

1. Repeated keys

Every field in a semantic layer YAML file repeats the same keys over and over (name:, type:, description:, etc.). A Graphene SQL field is one line: field_name: expression -- description #annotations. There’s no key to repeat, because position in the file is the schema, the same way amount FLOAT doesn’t need a type: FLOAT key beside it in a CREATE TABLE statement.

YAML’s object-list structure is built for arbitrary nesting; a semantic model doesn’t need arbitrary nesting. It needs a flat list of typed, named expressions, and SQL already has a compact, well-understood way to say that.

2. YAML doesn’t understand the SQL it’s holding

To a YAML parser, expr: "sum(case when revenue_recognized then amount else 0 end)" is an opaque string. It has no way to know that expression is an aggregation, so agg: sum has to be stated separately, even though the SQL already says so.

Graphene SQL doesn’t have this problem, because it owns its own SQL dialect end-to-end. The same parser that runs the query also reads the expression when the model is defined. It knows sum(...) makes a field a measure with no separate agg key needed, and for a handful of shapes it goes further: extract(year from dt) attaches timeGrain=year metadata automatically, with no annotation at all.

This matters now

Before AI, semantic layers were written by small data teams and never read directly by analyst users. A file’s real consumer was a BI tool’s frontend, which didn’t care about verbosity.

Agents care though. A 2-3x less compact semantic layer is 2-3x less room for what actually matters: the conversation; the dashboards it’s editing; the other tables it should have seen.

We think it’s time to move on from YAML. We hope you agree.


Want to see how good a state-of-the-art model is when armed with Graphene SQL? Drop us a line here.