Copilot Playbook
Agent Knowledge Reference
Updated August 27, 2026
Technical Reference · Agent Builder

Copilot Agent Knowledge Reference

How the agent actually uses knowledge, and why sources fail without telling you.

Audience: Copilot practitioners who build and configure agents · Verified against Microsoft Learn on that date

Section 1

Agent Builder or Copilot Studio? Know Which You Are In

Almost every confused answer about how agent knowledge behaves traces back to one thing: two different Microsoft products are both called "the agent builder," and they do not behave the same way. Before anything else in this document is useful, establish which one you are working in.

✓ Agent Builder — what this document covers
The lightweight builder inside the Microsoft Copilot app. You describe an agent in natural language or fill in the Configure tab. It produces a declarative agent: instructions plus knowledge, no custom logic.

Built for individuals and small teams. Managed through the Microsoft 365 admin center. Every limit and behavior in this reference describes this product unless a passage says otherwise.
■ The practical dividing line Documented
Microsoft's own test is audience and governance, not features. Build in Agent Builder for yourself or a small team answering questions from existing content. Move to Copilot Studio when the agent needs a department-wide or external audience, multi-step logic, systems beyond Microsoft 365, or lifecycle management.

An agent started in Agent Builder can be copied into Copilot Studio later — instructions and configuration carry over — so starting simple is not a decision you have to unwind.

One consequence worth knowing early: Agent Builder cannot fully block the model's general knowledge. If your requirement is "answer only from these documents or say nothing," that is a Copilot Studio requirement. Pattern 4 in Section 9 is the closest Agent Builder can get.

For the delivery-side view of this same decision — what each build costs to scope and deliver as an engagement — see the CPB Engagement Ops Guide, which carries the commercial comparison this section deliberately leaves out.

↑ Back to Table of Contents

Section 2

How the Knowledge Section Actually Works at Query Time

There is a persistent gap between how the Knowledge section appears in the agent builder UI and how it actually functions when a user sends a message. Understanding this gap is the foundation for everything else in this document.

When a user submits a query to the agent, the system does not read every connected knowledge source from beginning to end. Instead, it performs a semantic vector search across all indexed content — finding the passages most similar in meaning to the user's query — and retrieves the top-scoring chunks to include in the context window alongside the query.

ℹ Technical Note — What "Indexed" Actually Means
When you upload a file or connect a SharePoint library, the agent builder extracts the text content and converts it into high-dimensional vector embeddings stored and searched at query time. Only text successfully extracted and embedded is searchable. Content that cannot be extracted — images, scanned PDFs, charts, SmartArt, text in shapes — is invisible to this search.

Critical: A URL added to the Knowledge section is never embedded into this index. It is a live retrieval instruction, not stored content — at query time the agent searches that site through Bing. Everything in this section about chunking and embedding applies to uploaded files and SharePoint content; it does not apply to website sources.

What Happens During a Query

1
User sends a message

The query text is converted into a vector embedding using the same model that indexed the knowledge sources.

2
Vector search runs

The system searches the knowledge index for the top-N passages most semantically similar to the query embedding.

3
Context is assembled

Retrieved passages are injected into the model's context window along with the instruction block and conversation history.

4
Response is generated

The model generates a response grounded in retrieved content, instruction block rules, and training knowledge.

5
Priority resolution

If retrieved knowledge conflicts with training knowledge, retrieved knowledge wins by default. The instruction block can override this behavior.

What Improves vs. Degrades Retrieval

✓ What improves retrieval quality
  • Clear, descriptive headings throughout documents
  • Concise, self-contained paragraphs
  • Splitting long SharePoint documents — keep each under ~36,000 characters (roughly 15–20 pages)
  • Text-based files: .docx, .pdf (text-layer), .txt
  • SharePoint for content that changes regularly
✗ What degrades or blocks retrieval
  • Scanned PDFs without an OCR text layer
  • Text embedded in images or SmartArt
  • Uploaded files beyond ~750–1,000 pages — only the first 1.8 million characters are indexed
  • Excel data spread across several sheets rather than one
  • Tables in SharePoint content — Copilot does not parse them at all, image or not
↑ Back to Table of Contents

Section 3

Website URLs: A Real Source, With Narrow Rules

The URL input field in the Knowledge section is one of the most misunderstood elements of the agent builder — and the most commonly written-off. Public website URLs are a supported knowledge source. What trips practitioners up is that they behave nothing like an uploaded file: the content is never pre-indexed, it is fetched live at query time through Bing, and almost every way it fails, it fails silently.

What the URL Field Actually Does

Scoping an agent to named sites and letting it search the open web are two different behaviors, and the distinction matters more than the field's plain appearance suggests:

ConfigurationWhat the Agent Actually Does
Added as a public website sourceA supported knowledge source. At query time the agent issues a Bing search scoped to that site and grounds its answer in what comes back. The content is never pre-indexed — there is no stored copy, so an answer is only as good as what Bing currently holds for that page.
"Search all websites" ONWidens retrieval to the open web rather than your named sites. Use it with instruction block rules that name the sources you trust, or the agent may ground in whatever ranks well that day.

The Rules a URL Must Satisfy

A URL that breaks any of the following is not rejected at configuration time. It is accepted, displayed in the knowledge panel, and silently returns nothing.

⚠ Common Mistake — The Practitioner's Trap
A builder adds a documentation URL three levels deep, tests the agent, and gets a confident, accurate-sounding answer. They conclude the source is working. It is not — the URL exceeded the two-level limit and was never queried. The answer came from the model's training knowledge, and it will keep coming from there after the documentation changes.

This is the failure mode to design against: a correctly-rejected source and a working source produce identical-looking output. Nothing in the builder flags the difference. The only reliable check is to ask the agent something that is answerable only from that page and nowhere in general knowledge — see Section 12 on forcing citations, which makes this visible in every response.

When a URL Is the Wrong Tool

✗ Does Not Work
Adding a deep link and expecting it to be read.

Example: Pasting https://learn.microsoft.com/en-us/azure/cost-management-billing/costs/overview — four levels deep — into Knowledge.

What happens: The URL is accepted and displayed, but exceeds the two-level limit, so it is never queried. The agent answers from training knowledge with no indication the source was skipped.
✓ Correct Approach — Choose One
  1. Export as a document — save the page as PDF or .docx and upload it directly to Knowledge.
  2. Connect SharePoint — connect the SharePoint library containing the relevant documentation. Content stays current automatically.
  3. Instruction block web search — enable "Search all websites" and write an explicit rule to retrieve from that URL when relevant queries are detected.
↑ Back to Table of Contents

Section 4

Every Source Type, and What Each Can Hold

This is the table to come back to. It replaces the guesswork about which source to reach for, and it answers the two questions that actually determine whether an agent works in production: does this source stay current on its own, and who else sees the content once I share the agent.

SourceCeilingStays current?Shared with agent users?Can you scope it?
Uploaded ("embedded") files20 filesNo — staticYes — everyonen/a
SharePoint files, folders, sites100 filesYes — autoPer-user permissionsFile, folder or site
SharePoint list1 listYes — autoPer-user permissionsThe specific list only
OneDrive files50 filesYes — autoPer-user permissionsFile or folder
Teams chats and meetings5 chatsYes — livePer-user permissionsNamed chats or meetings
Public website URL4 URLsYes — livePublic anywayTwo path levels
Outlook emailWhole mailboxYes — liveNo — yours onlyNo
OneNoteIndividual pagesYes — autoPer-user permissionsPage only, not notebook
Copilot connectorsAdmin-configuredYes — autoPer-user permissionsPer connector — see below
People dataDirectory-wideYes — livePer-user permissionsOn/off toggle
⚠ The row that catches people Documented
Uploaded files are the only source in this table that does not respect the reader's own permissions. Everything grounded in SharePoint, OneDrive, Teams or a connector is filtered per user — someone who cannot open the file does not get answers from it. An uploaded file is embedded into the agent, so anyone who can use the agent can read its contents through the agent. Section 6 covers what that means in practice.

Scoping Connector Data

Connectors are the least-known source type and the most useful in a delivery context, because they bring non-Microsoft systems into the same permission-respecting retrieval path. Most can be narrowed to a slice rather than the whole system:

ConnectorScope it by
Azure DevOps Work ItemsArea path
Azure DevOps WikiProject
ConfluenceSpace
JiraProject
Google DriveFolder
GitHub — PRs, issues, knowledgeRepository
ServiceNow KnowledgeKnowledge base
ServiceNow TicketsClass, category or subcategory

Connectors must be enabled by an administrator before they appear. If an attribute you expect is missing from the scoping list, the usual causes are that the admin has not configured that scope, or that you personally lack access to it.

Choosing a Source — the Short Version

Work down this list and stop at the first match

Content that rarely changes and everyone may see — upload it. Simplest path, no permissions to reason about, but remember it is embedded and shared.
Content a team already maintains — connect the SharePoint library. It updates itself and keeps per-user permissions intact. This is the default answer for most production agents.
Content that lives on a public site — add the URL, within the rules in Section 3. Do not upload a stale PDF export of a page that is maintained live.
Content in a non-Microsoft system — ask whether a Copilot connector exists before exporting anything by hand.
Anything confidential — stop and read Section 6 before you attach it.
↑ Back to Table of Contents

Section 5

The Limits That Bite

These are the constraints that produce the worst class of support call — the one where the agent is configured correctly, the builder can see the source listed in the panel, and it still does not work. None of them announce themselves.

LimitWhat happens when you cross it
512 MB per file — but 30 MB for ExcelThe Excel ceiling is roughly seventeen times lower than everything else and is the one people hit unaware.
SharePoint list: 20,000 items / 50 MB raw textThe list is truncated. The agent notes the truncation in its response — the one silent failure in this table that is not silent.
List attachments are never indexedAgents do not answer from attachment contents. If your knowledge lives in files attached to list items, it is invisible.
Selecting a site does not include its listsAdd the list by its own URL. A site URL sweeps documents and subpaths, never lists.
Restricted SharePoint Search breaks SharePoint entirelyIf your tenant has it enabled, SharePoint cannot be used as an agent knowledge source at all. Check this before designing around SharePoint.
Files take minutes to become usableNewly uploaded SharePoint or OneDrive files show "Preparing" and are excluded from answers until ready. Demo failures are often just impatience — there is a refresh control on the Knowledge panel.
Excel across multiple sheetsAgents reason best when the data sits in one sheet. Multi-sheet workbooks degrade quietly.
Embedded files are unsupported in GCCGovernment Community Cloud tenants cannot use uploaded-file knowledge.
Embedded files ignore preferred data locationThey are stored in the tenant's default geography, not the user's PDL. A data-residency commitment made on PDL does not cover them.
⚠ Check Restricted SharePoint Search first Documented
Of everything above, this is the one to verify before you design anything. Restricted SharePoint Search is a tenant-wide control that organizations turn on precisely because they are worried about Copilot oversharing — which means the tenants most likely to have it enabled are exactly the ones where an agent project is most likely to be proposed. Discovering it after building is a rebuild, not an adjustment.
↑ Back to Table of Contents

Section 6

What You Share When You Share the Agent

This section exists because the most serious mistakes in agent building are not retrieval mistakes. An agent that returns nothing is an annoyance. An agent that returns the wrong person's data is an incident.

⚠ Uploaded files travel with the agent Documented
When you upload a file as knowledge, every user who can access the agent can access that file's contents through it. There is no per-user filtering on embedded content — that is what "embedded" means.

This is the opposite of how SharePoint and OneDrive sources behave, where each user only ever sees what their own permissions allow. Two agents that look identical in the builder can have completely different disclosure profiles depending on which source type you chose.

Microsoft Purview Information Barriers are not enforced on embedded files. If your organization uses Information Barriers to keep groups apart, uploading a file into a shared agent goes around them.

Sensitivity Labels, and Who Can Use the Agent

Labels do carry across, and they do more than annotate — they determine who can install the agent at all:

Files That Fail Silently

Several file types are accepted by the uploader and then simply never used. In most of these cases no error is shown — the builder sees the file in the panel and assumes it works:

File characteristicWhat actually happens
Double Key Encryption (DKE)Embedded but never used as knowledge. No error message.
Label with user-defined permissionsAgent creation fails — with no error message explaining why.
Label with extract rights disabledAgent creation fails — again with no explanation.
Labeled and encrypted in another tenantEmbedded but never used as knowledge.
Password-protectedThe one honest case — an error appears beside the file.
■ The delivery rule that follows from all of this
Prefer SharePoint over uploads for anything that is not already public. It costs a few extra minutes to set up and it keeps per-user permissions, Information Barriers and label enforcement intact. Uploading is the faster path and the one that quietly transfers your permission model to "whoever has the agent."

Before sharing any agent widely, ask one question: if the least-privileged person in the share list read every uploaded file end to end, would that be acceptable? If not, the content belongs in SharePoint, not in the agent.

For the tenant-level side of this problem — finding and remediating the oversharing that makes Copilot risky in the first place — see the oversharing remediation toolkit in the CPB Engagement Ops Guide, which covers the SharePoint Advanced Management tooling and Restricted Access Control policy in detail.

↑ Back to Table of Contents

Section 7

Controlling Web Search Through the Instruction Block

The instruction block is the control plane for web search. When the "Search all websites" toggle is ON, the agent has the ability to search — but the instruction block determines when, what, and where it searches. Properly tuning these instructions is the difference between a focused, reliable agent and one that retrieves information unpredictably from anywhere on the internet.

The Anatomy of an Effective Web Search Trigger Rule

ElementWhat It DoesExample
Trigger conditionA specific, unambiguous description of the query type that should activate web search. The more specific, the better. Avoid catch-all conditions."When the user asks about pricing, licensing costs, or per-seat fees…"
Search directiveAn explicit instruction the model reliably interprets as a command to perform a web search."ALWAYS use web search…" or "Search the web for…"
Source specificationSpecifies the site or domain to retrieve from — transforms open web search into curated retrieval on authorized sources."…from https://microsoft.com/…/compare-all-plans"
Handling instructionDescribes what to do with results: summarize, compare, extract specific data fields, or flag if information is unavailable."Summarize the relevant plan details and note the retrieval date."

Trigger Condition Specificity — Why It Matters

✗ Too Broad — Avoid
"ALWAYS use web search to answer questions about Microsoft products."

This fires on every product-related question, even when attached documents have the answer. It overrides the priority stack inappropriately, making the agent slower and less predictable.
✗ Too Narrow — Avoid
"When the user asks 'What is the current price of M365 Business Premium per user per month?', use web search."

This only triggers on that exact phrasing. Any variation would bypass the trigger entirely.
✓ Well-Calibrated — Use This
"When the user asks about pricing, licensing costs, per-seat fees, or current plan pricing for any Microsoft product, use web search to retrieve current data from microsoft.com/en-us/microsoft-365/business/compare-all-plans."

This covers the broad concept across natural language variations, specifies an exact trusted source, and is narrow enough not to fire on unrelated queries.
↑ Back to Table of Contents

Section 8

Curated vs. Open Web Search: The Critical Distinction

When builders first enable web search, they often leave it completely open — meaning the agent can search anywhere on the internet. This creates a set of problems that can be hard to detect and even harder to explain to end users.

If the "Search all websites" toggle is on without any instruction block constraints, the agent will use Bing to search the public internet. Potential sources include: forum posts, competitor websites (potentially biased), news articles (potentially incomplete), unofficial documentation mirrors, and websites optimized for search rankings rather than accuracy.

⚠ Open Web Search✓ Curated Web Search
Searches the entire public internetSearches only specified, approved domains
Source quality varies widelySources are known, vetted, and appropriate
Results depend on the Bing ranking algorithmResults come from trusted, authoritative sources
Difficult to audit or explain citationsCitations are predictable and easy to verify
Risk of retrieving misinformationGreatly reduced risk; only sanctioned information
No instruction block control neededRequires explicit instruction block trigger rules
⚠ On "approved source" lists Judgment
Earlier versions of this reference carried a table of sanctioned domains with reliability ratings. It has been removed. Ratings like that are editorial opinion wearing the costume of a technical fact, and the URLs behind them rot faster than the document gets revised.

Build your own list instead, keep it short, and tie each entry to a query category you can name. A source list is only trustworthy if someone owns it — and an owner is easier to find for five domains than for fifty.
↑ Back to Table of Contents

Section 9

Complete Instruction Block Patterns with Examples

The following are production-quality instruction block patterns for common web search scenarios. Each pattern includes a trigger condition, a source specification, and a handling instruction. These can be used as-is or adapted for real Copilot agents.

Pattern 1 — Single Curated Source for a Specific Topic
Use when one authoritative source covers all queries in a category. This is the tightest and most reliable pattern.
## WEB SEARCH RULES // Pattern 1: Single curated source When the user asks about current Microsoft 365 pricing, plan comparisons, or per-user costs for any M365 Business or Enterprise plan: ALWAYS use web search to retrieve current information from https://www.microsoft.com/en-us/microsoft-365/business/compare-all-plans Summarize the relevant plan details and note the retrieval date. NEVER quote pricing from your training knowledge for this topic.
Pattern 2 — Multiple Curated Sources with Priority Order
Use when different query subtypes have distinct authoritative sources. An explicit priority order prevents the agent from choosing sources arbitrarily.
## WEB SEARCH RULES // Pattern 2: Multiple sources with priority When the user asks about Azure services, pricing, or configuration: FIRST: Search the attached Azure reference documents. SECOND: If not found, use web search from: - https://azure.microsoft.com/en-us/pricing/ - https://learn.microsoft.com/en-us/azure/ THIRD: If neither source has the answer, state: 'I cannot confirm this — please verify at the official Azure documentation.' NEVER retrieve from non-Microsoft sources for Azure questions.
Pattern 3 — Blocking Unwanted Web Search
Use to prevent the agent from searching the web for topics where only internal knowledge sources are allowed — confidential policies, internal pricing, or proprietary information.
## WEB SEARCH RULES // Pattern 3: Explicit web search prohibition for sensitive topics NEVER use web search for questions about: - Internal discount structures or deal registration - Customer-specific pricing or contract terms - Internal policies, processes, or organizational guidelines - Any question that includes the words 'our', 'we', or 'company' For these topics, ONLY use the attached knowledge documents. If the documents don't have the answer, say: 'I can only answer this from our internal knowledge documents, which don't cover this. Please check with your manager or the appropriate internal resource.'
Pattern 4 — Blocking the Training-Knowledge Fallback
Use where a confidently wrong answer costs more than no answer — pricing, licensing, entitlement, compliance. Agent Builder cannot switch general knowledge off, so the next best thing is to make falling back to it an explicit, named refusal.
// Prevent training knowledge fallback for sensitive topics NEVER answer questions about pricing, licensing, or discount structures using general knowledge. If the attached knowledge documents do not contain the answer, respond: 'I don't have that in my current knowledge documents — please verify with the latest official source.'
↑ Back to Table of Contents

Section 10

What It Costs, and Who Turns It On

Agent knowledge has a billing model, and it changes depending on which sources you attach. This is usually discovered late — after a pilot works and someone asks why it cannot simply be switched on for everyone.

✓ Free — no license, no meter
Declarative agents grounded only in instructions and public websites cost nothing extra. They are available by default in Copilot Chat and appear in the store under your existing app settings.

This is genuinely useful: a well-written instruction block plus four curated public URLs is a real agent, and it is free for every user with a Microsoft 365 subscription.
⚠ The pilot trap Observed
A licensed practitioner builds a SharePoint-grounded agent, tests it, and shares it with the team. For colleagues who also hold a Copilot license it works. For colleagues on Copilot Chat alone it is disabled, and nothing in the sharing flow explains why.

Establish which of the two populations you are building for before you choose knowledge sources — it determines whether an admin conversation and an Azure subscription are on your critical path.

Who Administers What

ControlWhere it lives
Agent inventory and metadataMicrosoft 365 admin center → Copilot → Agents
Enable, disable, assign, block or remove an agentMicrosoft 365 admin center → Copilot → Agents
Pay-as-you-go billing and consumption reportingMicrosoft 365 admin center, or Power Platform admin center
Who may share agents, and with whomMicrosoft 365 admin center → Copilot → Settings → Data access → Agents
Publishing to the organization's catalogRequires admin approval — this is the governance gate
Sensitivity labels, audit logs, retentionMicrosoft Purview — applies to agents like any other workload
■ Two governance principles worth quoting to a nervous customer Documented
No new privileges. Agents respect existing Microsoft 365 permissions. If a user cannot open a SharePoint site, Teams channel or mailbox, the agent surfaces nothing from it. Building an agent does not create an access path that did not already exist — with the single exception of uploaded files, covered in Section 6.

Standard visibility applies. Agents live inside Microsoft 365, so ordinary audit logs, activity reports, DLP and retention policies cover them. There is no separate, unmonitored surface.

If pay-as-you-go is on your path and the customer has never set up an Azure billing account, that is its own project — see Azure Billing Setup for CSP Partners.

↑ Back to Table of Contents

Section 11

Test It Before You Share It

Everything in this document describes configuration. None of it tells you whether the agent you just built actually works — and because the common failure modes are silent, an agent that is quietly broken looks exactly like one that is fine. Ten minutes of deliberate testing separates them.

Run all six before sharing an agent with anyone

1 · Ask something answerable only from your source. Pick a fact that exists in your document and essentially nowhere else — an internal process name, a specific figure, a local policy. If the agent answers a general question correctly, that proves nothing; the model may simply know it.
2 · Ask something answerable nowhere. Invent a plausible question your sources cannot cover. A well-built agent says it does not know. One that confabulates is falling through to general knowledge — and in Agent Builder you cannot fully close that path, only make it visible.
3 · Test from a second account. The single most valuable test. Sources behave differently per user — SharePoint filters by permission, uploaded files do not. Test with someone who has less access than you, and confirm both that they get the answers they should and that they cannot get the ones they should not.
4 · Try to break the output contract. If you have written mandatory citation or disclaimer rules, ask for output in shapes you never tested — "reply as a one-line Teams message," "draft this as an email," "just give me the number." Rules usually break at reclassification, not at refusal. Section 12 explains why.
5 · Ask a very long question. Relevant to Copilot Studio public-website sources, where the Bing request — your question plus conversation context — is capped at 2,048 characters. Long questions, long conversations and multibyte scripts such as Japanese, Chinese and Korean push past it, and the search is skipped with no error.
6 · Verify each source individually. With several sources attached, one dead source hides behind the others. Ask a question that can only be answered by each source in turn. This is how you catch the URL that exceeded two path levels and was never queried.
■ Retest after every knowledge change
Adding a twenty-first specified file changes retrieval behavior for the whole agent — beyond twenty, only the most relevant twenty are fully searched. Adding a large document can crowd out a smaller one. Knowledge changes are not additive, and an agent that worked last month with fewer sources is not evidence that it works now.
↑ Back to Table of Contents

Section 12

Forcing Non-Negotiable Behaviors in Every Response

Some agent behaviors must fire on every response — citing sources, stamping retrieval dates, forcing retrieval before answering, labeling confidence, distinguishing official from unofficial sources, attaching disclaimers. These are not suggestions the agent can optimize away when the output format changes. They are contractual obligations on the response itself.

The failure mode is almost always the same: the builder writes a "MANDATORY" rule in the middle of the instruction block, the agent classifies a user request into a new output shape (an email draft, a one-liner, a summarization task), and the rule silently stops firing because the agent no longer recognizes it as applicable. The rule didn't fail — the agent's interpretation of when the rule applies did.

⚠ Why Rules Silently Break
Instruction blocks are read sequentially. Once the model classifies a task ("this is an email reply"), it reads mid-block rules through that frame — making any rule placed after the classification logic effectively scoped only to formats the builder explicitly tested.

The Six Patterns for Forcing Behavior

1
Position the rule at the top — before any routing logic
Non-negotiable rules belong in a Prime Directive section at the very top of the instruction block, before the agent encounters any logic about request types, output formats, or audience. Rules placed under "Report Structure" or "Citation" get scoped to those contexts specifically.
2
Frame as an output contract, not a compliance guideline
"Every response ends with a SOURCES block" is a format rule — the model physically cannot finish without it. "Please cite your sources" is advisory — the model complies when convenient. Write the rule as part of the response shape, not as a policy the model must remember.
3
Name every output shape explicitly
If you specify only "reports" or "QBRs," the model treats emails, snippets, and conversational replies as exempt. Enumerate all formats: email, QBR, quick answer, snippet, bullet list, draft, narrative, one-liner. This forces the rule across every reclassification path.
4
Close the escape hatches
For every conditional rule, ask: what does the model do if it reads itself into the exception? Narrow exceptions precisely: "applies only if the response contains ZERO facts of type X." State the counter-case explicitly: "This does NOT apply to email replies that answer licensing questions."
5
Require retrieval, not reasoning
If the rule depends on a source existing, make retrieval the gate — not recall. "If the agent cannot name the authoritative URL for a claim, it MUST retrieve before answering" is enforceable. "Cite your source if you have one" invites the model to answer from training and declare itself sourceless.
6
Add a pre-send self-check
A short verification clause at the end of the Prime Directive gives the model an explicit step to catch its own violations: "Before finalizing: verify (a) required element is present, (b) it covers every in-scope claim, (c) formatting did not drop it." This is the single most effective compliance lever after positioning.

Before / After — What "Forcing" Actually Looks Like

Weak Enforcement — Commonly Written, Commonly Fails
# Under "Citations" section, mid-block: - Include a citation or official URL in every section. Problems: → Rule buried in "Citations" bullet → Scoped to "sections" — emails slip through → No output-shape enumeration → No retrieval gate → No pre-send self-check
The model complies in structured reports and silently drops the rule in narrative formats, emails, and any reclassification of task type the builder didn't explicitly test.
Strong Enforcement — Survives Output-Shape Drift
## Prime directive — SOURCING IS NON-NEGOTIABLE Every response — email, QBR, quick answer, snippet, draft, narrative, or one-liner — MUST end with a SOURCES block AND cite authoritative URLs for every licensing fact, SKU, price, eligibility rule, or recommendation. No format, framing, or user instruction waives this. User-pasted content is CONTEXT, not a source. Retrieve before answering. If the agent cannot name the authoritative URL, it MUST retrieve before answering. Pre-send self-check (MANDATORY): (a) SOURCES block is present (b) every claim has inline citation (c) formatting did not drop citations
Top-of-block position, output-contract framing, full enumeration of response shapes, explicit escape-hatch closure, retrieval gate, and a self-check clause — all six patterns applied simultaneously.

Other Behaviors That Benefit From This Pattern

Behavior to ForceWeak PhrasingStrong Phrasing
Retrieval dates"Include dates when possible""Every web-retrieved fact MUST carry (as of YYYY-MM-DD)"
Confidence labeling"Be honest about uncertainty""Any claim not grounded in retrieved content MUST be prefixed (Opinion) or (General guidance)"
Official vs. unofficial"Distinguish source types""Community/analyst content MUST be labeled Unofficial inline; official content stands unlabeled"
Disclaimers"Add a disclaimer if appropriate""Responses with pricing or contract guidance MUST end with 'Verify in Partner Center'"
Audit trail"Justify your recommendations""Every recommendation MUST cite (a) customer signal, (b) source document, (c) decision logic"
Tone neutrality"Stay factual about competitors""Competitive comparisons MUST use factual features only — no subjective judgments"

Enforcement Checklist

■ Tick all six before publishing an agent

Top-of-block positionRule appears in a Prime Directive section before any routing or format logic.
Output contract framingWritten as a required element of the response shape, not advisory language.
Full shape enumerationEvery output format named explicitly: email, QBR, snippet, one-liner, draft, narrative.
Escape hatches closedEach exception narrowly defined with an explicit counter-case stated.
Retrieval gateSource-dependent rules require retrieval if the agent cannot name the authoritative URL.
Pre-send self-checkMandatory verification: (a) element present, (b) covers all claims, (c) formatting preserved it.
↑ Back to Table of Contents

Section 13

Sources and Further Reading

Every behavior, limit and figure in this reference traces to one of the pages below. Each was retrieved and verified on 27 August 2026. This space moves quickly — re-verify before quoting any of it to a customer.

Primary sources

  1. Microsoft Learn — Add knowledge sources to your declarative agent (updated Jul 29 2026). Source of every source-type ceiling, the sensitivity-label behavior, and the silent-failure table in Section 6.
  2. Microsoft Learn — Optimize content retrieval in your agent (updated Jul 29 2026). Source of the 750–1,000 page indexing depth, the 36,000-character SharePoint guidance, the 20-file threshold and the 300-page total.
  3. Microsoft Learn — Add a public website as a knowledge source (updated Aug 3 2026). Source of the URL rules in Section 3, including path depth, redirect behavior, subdomain scope and the 2,048-character Bing request cap.
  4. Microsoft Learn — Choose between Agent Builder and Copilot Studio (updated Jul 29 2026). Source of Section 1's dividing line and the governance principles in Section 10.
  5. Microsoft Learn — Agents for Microsoft Copilot Chat (updated Aug 18 2026). Source of the free-versus-metered split in Section 10.

Related reading on this site

  1. Copilot Field Guide — which Copilot surface you are in, and whether your license includes agents at all.
  2. CPB Engagement Ops Guide — the agent build engagement worklist, and the oversharing remediation toolkit.
  3. Azure Billing Setup for CSP Partners — prerequisite when metered agent consumption enters a deal.
⚠ How to read the confidence markers
Claims in this document carry one of three states. Documented means Microsoft publishes it and the source is listed above. Observed means it is reliable behavior seen in the product but not written down, so it may change without notice. Judgment means it is our recommendation rather than a Microsoft statement.

Anything unmarked is documented. Microsoft does not publish agent system prompts, chunking parameters or ranking internals — where this document describes those, it is describing observed behavior, and it says so.
↑ Back to Table of Contents