Field note · AI and the semantic layer
An agent needs to know what your data means, how it fits together, and when to trust it. Choosing a model is only part of that work.
01 · Above and below
A launch video shows the agent answering a question in plain English. It does not show which table the answer came from, which rows were excluded, or who decided that a canceled trial still counts as a customer. You need those answers before you can judge whether to trust the result.
The drawing exaggerates the problem, but the distinction is useful: a working demo tells you little about the data supporting it.
02 · The rented part
Choosing and evaluating a gen AI model matters. But access to that model does not give it the knowledge your team has built up about the business.
Consider a definition of active subscriber that survived three arguments with finance, a join graph checked for fan-out, and a record of which tables are safe to trust before 9am. Making that knowledge available to an agent takes work within your own team.
An agent can carry a mistaken definition through several tool calls. A later step may look reasonable even though it depends on an earlier query that counted the wrong customers.
03 · The context checklist
For an agent querying a warehouse, I would check the following. Having staging, intermediate, and mart models helps, but the agent still needs to know how to use them.
Set limits on what the agent can do as well: read-only credentials, row caps, a fixed set of exposed topics, and no path that bypasses the layer. These controls limit access and resource use; they do not establish that an answer is correct.
A model upgrade will not settle your subscriber definition or assign someone to investigate a failed data check.
04 · Warehouse maintenance
In an undocumented warehouse, two tables can look safe to join even when they represent different things. Analysts run into this too. An agent needs enough information to distinguish a valid join from one that merely runs.
The usual analytics engineering work helps here: tests that fail when a key stops being unique, consistent naming and casting in staging, and fact and dimension models with a stated grain.
Keep the reasons for those choices too: why revenue excludes credits, why the trial table is snapshotted nightly, and which trade-off the team accepted. Write them down while the reasoning is still fresh.
Prompt instructions and retrieval can help an agent find this information. Someone still has to maintain it.
05 · The new failure mode
A broken join can go unnoticed in a dashboard. Sometimes a colleague catches it because they know roughly what the number should be.
An agent can turn that same number into a fluent recommendation to shift a budget. If the reader cannot inspect the query or recognize an implausible result, the explanation may make the error more persuasive.
Where a workflow removes human review, more of the checking needs to happen before the answer is delivered. Shared definitions and tested models become especially useful there.
06 · The translation problem
We routinely leave details unstated when talking to colleagues. That works when we share the same assumptions.
“How many active customers do we have in France?” is a perfectly good English sentence and an underspecified query. Active when, by which signal? A customer at the account level or the seat level? France by billing address, by IP, or by the sales territory that owns the account? A colleague may know the convention, or know whom to ask.
Without those definitions, the same question can produce two defensible queries with different results. The reader needs to know which interpretation was used.
The semantic layer records agreed definitions so the agent can reuse them. It still needs to choose the right metric and filters, and ask for clarification when the question is ambiguous.
07 · The quiet side effect
An experienced analyst can compensate for gaps in a model because they know the business. That knowledge is difficult to share with a new colleague, let alone an agent, unless it is recorded.
Fields such as ai_context give teams a place to record
guidance alongside model definitions. Use them to explain exclusions,
preferred join paths, and questions a metric should not answer.
Include that guidance in review. A technically valid definition can still leave its intended use unclear.
I use the established term semantic layer, but prefer data context layer. It better describes what I want from it: enough context for a person or an agent to interpret a number correctly.
My proposed ownership split is for analysts to define meaning and engineers to validate the implementation, with both reviewing before publication. I explain the reasoning in the semantic layer ownership model.
08 · From prototype to production
A prototype gives stakeholders something concrete to react to and helps uncover disagreements about definitions. Keep it in a sandbox while you work those out, with a separate review before production use.
The job
Make the data work part of the plan.
When planning an agent demo, include time to review the definitions, joins, permissions, and wrong answers it exposes. That work will also help the people already using the warehouse.