Grounding Copilot agents properly
The question sounds harmless: “Which category causes the most negative feedback?” An agent on the feedback table answers it immediately, fluently and with numbers. The problem is not that it refuses. The problem is that it answers even when the piece of feedback that actually matters does not appear in its data at all.
Grounding is usually discussed as a property of the agent: better instructions, a narrower role, a sentence against hallucination. The more effective part sits one layer down. It happens in the source, long before anyone asks a question.
This article walks through it on a real table: Kundenfeedback in Dataverse,
four records, two AI-backed columns. Everything here was checked in a live
environment — not retold from the docs.
What you see first
The table has four custom columns: Name as the primary column, Feedbacktext
as multi-line free text, plus Kategorie and Sentiment. Behind those sit 22
system columns.

Three rows carry a category, one does not. And of all rows, the empty one is the interesting one:
Lieferung Bestellung 40021
"The goods arrived four days late and the box was dented.
Support did not reply to two emails."
Kategorie: — Sentiment: —
That is the only unambiguously negative piece of feedback in the table. For an
agent that evaluates through Kategorie and Sentiment, it does not exist. The
agent will not say so. It will answer about the other three, and that answer
will sound right.
What actually happens technically
Kategorie and Sentiment look like text columns in the column list — the
column definition says otherwise:

Data type Prompt, format Text. The value does not come from typing, but
from an instruction bound to exactly one input column:
Klassifiziere das folgende Kundenfeedback in genau eine Kategorie:
Lieferung, Produktqualität, Preis, Service, Sonstiges.
Gib nur das Kategoriewort zurück.
Feedback: ⟨ Kundenfeedback.Feedbacktext ⟩
And for Sentiment:
Bestimme das Sentiment des folgenden Kundenfeedbacks
(positiv, neutral oder negativ) und begründe deine Einschätzung
in genau einem Satz.
Feedback: ⟨ Kundenfeedback.Feedbacktext ⟩
That already contains the actual point: the structure an agent relies on is created at write time — not at answer time. Classification happens once, when the record is created or the input column changes. The agent only reads it back later.
The note at the top of the panel is not fine print:
Die Werte der Promptspalte werden asynchron generiert. (The values of the prompt column are generated asynchronously.)
How these columns work in detail — triggers, status codes, cost — is covered in Prompt columns in Dataverse. What matters here is only what they mean for grounding an agent.
The obvious shortcut that isn’t one
“Then I’ll just let the agent read the free text.” It sounds cheaper and is the more expensive route. Classifying free text at answer time gives you a new judgement on every question — not reproducible, not filterable, not auditable. Two identical questions can produce two different numbers, and nobody can show afterwards where either came from. Aggregates over free text are not analysis, they are an estimate with decimal places.
“Then I’ll write the agent stricter instructions.” Instructions set the tone and the scope. They do not set what is in the table. No system prompt helps against an empty column.
What holds
The allowed vocabulary is the real guardrail
The invoice feedback — a line item that was never ordered — sits in the table as
Service. Not as “Rechnung” (invoice). That is not a model error, it is the
honest answer within the vocabulary the column prescribes: Lieferung,
Produktqualität, Preis, Service, Sonstiges. A category “Rechnung” does not
exist there.
From that follows an uncomfortable division of labour: if you want to evaluate invoice topics, you change the column, not the question you ask the agent. The vocabulary of the source is the limit of what the agent can ever distinguish.
A forced justification makes the derivation inspectable
The sentiment prompt demands one sentence of reasoning. It costs space and is worth every character:
"Das Sentiment des Kundenfeedbacks ist positiv, da trotz eines anfänglichen
Problems mit einer falschen Position auf der Rechnung die schnelle und
unkomplizierte Korrektur innerhalb eines Tages hervorgehoben wird."
You can argue with that “positiv” — it was a billing error after all. That is
exactly why the justification is valuable: it makes the rating contestable. A
bare positiv would have gone unchallenged.

The agent inherits the user’s permissions
The boundary that holds most reliably is the one that does not live in the agent. An agent on a Dataverse table sees what the asking person is allowed to see — security roles, table permissions and business units still apply. That is exactly why trying to govern access through the agent’s identity is a dead end; Entra Agent ID explains why.
Views are the curated surface
A table with 26 columns is a poor basis for answers. A view holding exactly the four relevant columns and a sensible filter is a good one. Whatever you keep out of the view, the agent never has to interpret.
A checklist before your first agent
- Are there structured columns, or only free text? Without structure an agent can only paraphrase, not evaluate.
- Do you know the vocabulary? The column’s category list is the limit of every later evaluation. Read it before asking a question it cannot answer.
- Are the values complete? Records that existed before the column stay empty. There is no backfill over existing data.
- Do you check the status, not just the value? The values are generated asynchronously. Right after saving, the column is empty — that is not a fault, that is the normal case.
- Do you force a justification? A rating without reasoning cannot be checked.
- Are the permissions right? The agent shows what the asking person may see. Test with an account without admin rights, not with yours.
- Is the switch even on? The column definition carries “Ausführung der Spalte ‘Prompt’ zulassen”. With it off, everything stays empty — without an error message.
What stays uncomfortable
No backfill over existing data. The oldest row in my table has neither category nor sentiment, and will not get them as long as its feedback text stays unchanged. Anyone introducing such a column later has a blind spot across all historical records — and the agent says nothing about it.
The value reflects the input as of its run. I created a row whose feedback text was still empty at that moment. The column duly judged “no substantive statements … therefore neutral”. After the text was filled in, that assessment stood for a while before the rerun corrected it. Ask inside that window and you get a clean, well-reasoned, outdated answer.
Generative does not mean deterministic. The same feedback can be classified differently on a rerun. For a trend that is fine. For a metric that triggers an escalation, it is not.
Structure is work somebody has to maintain. Vocabularies age, new topics appear, the category list becomes a maintenance item. That is the honest price for answers that stay verifiable.
See also
- Prompt columns in Dataverse — how the columns that create this structure actually work.
- Entra Agent ID — why identity is not permission.
- UI components in chat — when the agent does not just answer but offers to act: the same principle, one step further.