
Your internal AI gives a precise answer from a document the employee could not open in Google Drive.
The answer may be factually correct and still represent an access-control failure.
Permission-aware company search is not achieved by hiding the source link after retrieval. The system must know who is asking, which source permissions apply now, which content may enter the model context, and what must happen after access, classification or retention changes.
This part owns retrieval permissions: who may search which content, and what must change when access does. Keeping the index itself current (document identity, updates and deletion) is part 5; deciding what the assistant may do with what it finds is part 7. The full build path is on the guide page.
The answer in 60 seconds
A secure company knowledge layer needs one permission chain:
- authenticate the person or service making the request;
- preserve the source document’s identity, classification and access reference;
- filter candidates for that principal;
- authorise sensitive content before it reaches the model;
- cite only sources the principal can open;
- propagate revocation, deletion and retention changes;
- log enough to investigate without creating another uncontrolled copy.
If your index uses one broad service account and cannot reproduce why a user was allowed to retrieve a document, it is not permission-aware yet.
In this article
- Start with a principal, not a chat window
- Choose the right source access model
- Carry permissions and classification into the knowledge layer
- Authorise before content reaches the model
- Propagate revocation, deletion and retention
- Test and operate the permission chain
1. Start with a principal, not a chat window
Every retrieval request needs a named principal: a user, group-aware delegated identity or bounded service identity. “The company AI asked” is not enough.
This matters because connected systems use different access models:
- Google Drive evaluates a file or folder’s access-control list and the current user’s capabilities;
- Slack scopes define what an app or user token may request, while the Conversations API filters conversation types by token scope;
- GitHub App permissions combine app permissions with installation, repository and sometimes user access;
- Linear OAuth separates read, broad write and narrower create permissions.
Authentication answers “who is this?” Technical permission answers “what can this identity reach?” Your business design still has to answer “may this content be used for this purpose?”
Treat the source as the authority for access whenever practical. A copied index is a retrieval aid, not a new owner of confidentiality.
2. Choose the right source access model
There are two common integration models, and they have different failure modes.
Delegated access retrieves on behalf of the requesting user. It naturally reflects much of that user’s current source access, but token lifecycle, group changes and per-source semantics still need handling.
Service access uses an integration identity. It is simpler for scheduled ingestion, but the index can inherit everything that service account can see. If that identity crosses HR, finance, commercial and delivery domains, the retrieval layer may flatten boundaries that the source systems maintained.
Notion makes the distinction concrete. Its authorization guide says an internal connection must be explicitly shared with pages before it can access them. The broader connection overview distinguishes internal connections, public OAuth connections and personal tokens by identity and content access. A bot that can read a page is not evidence that every user of your assistant may read it.
For each connector, record:
- integration identity and owner;
- granted scopes and accessible domains;
- whether requests are delegated or service-based;
- group and object-level checks;
- token expiry, rotation and revocation;
- the process for role changes and offboarding.
3. Carry permissions and classification into the knowledge layer
A vector or search index should not reduce a document to text plus an embedding. Preserve enough control metadata to make a retrieval decision:
source_system, source_id, source_url, owner, classification, acl_reference, version, modified_at, retention_state, deleted_at.
Do not treat an embedding as harmless technical metadata. OWASP’s Vector and Embedding Weaknesses identifies unauthorised access, cross-context leakage, poisoning and weak logging as RAG risks. Its Sensitive Information Disclosure category includes personal, financial, legal and proprietary information.
Use partitions and metadata filters to reduce the candidate set, but understand their limits. An ACL snapshot can become stale between source updates. A group name can change meaning. A document moved to a confidential folder may keep an old index label until the next sync.
Google recommends the narrowest practical Drive API scope; a narrow OAuth scope still does not replace file-level permission checks. In the automation layer, n8n RBAC and project roles and workflow sharing can bound who works with workflows and credentials. They do not automatically reproduce every source ACL inside a vector store.
4. Authorise before content reaches the model
The safest rule is simple: if the principal may not see a document, its text must not enter the prompt or model context.
Post-generation redaction is too late. The model has already received the information and may reveal it indirectly through an answer, summary, comparison or refusal.
Use a layered decision:
| Layer | Question | Fail-closed behaviour |
|---|---|---|
| Identity | Who is asking, in which tenant and role? | Reject anonymous or ambiguous principal |
| Candidate filter | Which domains and classifications are eligible? | Exclude unlabelled or mismatched content |
| Source authorisation | Can this principal access these exact objects now? | Remove unauthorised chunks before context |
| Generation | Can the answer cite only authorised evidence? | Abstain when evidence is missing |
| Presentation | Can the user open every citation? | Hide the answer, not merely the link |
| Audit | Can we reconstruct the decision? | Record a bounded denial/failure event |
For sensitive domains, check authorisation against the source of truth at retrieval time. The AWS Security Blog’s RAG authorisation pattern explains why periodically synced metadata can lag source permission changes and demonstrates source-backed authorisation before chunks reach the model. It is one vendor pattern, not the only valid architecture.
5. Propagate revocation, deletion and retention
Permission-aware search has to remain correct after launch.
Google Drive’s change logs expose current state changes for users and shared drives, but complete coverage requires the relevant logs and a change entry is not a property-level delta. Other sources use webhooks, polling, events or scheduled reconciliation. Your design needs an explicit answer for:
- access revoked from a person or group;
- a confidential page moved or re-shared;
- a user leaving the company;
- a source document deleted or replaced;
- a retention period ending;
- a connector token losing access;
- an index update failing halfway.
Deletion must cover raw copies, parsed text, chunks, embeddings, caches and generated material where applicable. Retention must be attached to purpose and source class, not set to “forever because search may need it”.
Keep credential management separate from document permission logic. n8n’s supported external secret stores can improve how eligible connector credentials are resolved. They do not decide which employee may retrieve which document.
Related: IZZY’s guide to managing AI agent permissions.
6. Test and operate the permission chain
Test denials as seriously as useful answers.
Build a small matrix with real role patterns and synthetic or authorised test documents:
- allowed user retrieves and opens the expected source;
- disallowed user retrieves nothing from that domain;
- group removal closes access;
- document reclassification changes retrieval;
- deletion removes every derived representation;
- stale or failed ACL sync causes abstention, not permissive fallback;
- logs contain the decision without reproducing unnecessary sensitive text.
NIST’s AI Risk Management Framework treats testing, use and evaluation as part of ongoing risk management. In n8n, a security audit can surface listed credential, node, filesystem and instance risks; external log streaming can send selected events to an existing monitoring process on eligible plans. Neither proves that your end-to-end permission chain is correct. Test it.
Illustrative Atlas access change
Atlas has a confidential acquisition workspace. Léa can search it while she is on the deal team. When she transfers to delivery, the source removes her group access.
A sound company brain receives or detects the change, removes acquisition content from her eligible retrieval set, invalidates relevant caches and confirms that an acquisition question now abstains. A copied index that continues answering until the next weekly rebuild has failed, even if its answer is accurate.
Atlas is illustrative, not an IZZY client case. No system or result was tested.
Conclusion: preserve the source boundary
A company brain should make authorised knowledge easier to use, not make confidential knowledge easier to bypass.
Name the principal. Preserve source identity and classification. Authorise before model context. Propagate revocation, deletion and retention. Test the negative path. Keep the source system as the authority for access.
Review one permission-sensitive domain with IZZY
Bring one domain such as HR, finance, commercial proposals or client delivery, plus its source system, user groups, retention rule and one access-change example. We will map the permission chain and identify where the design should delegate, filter, re-authorise, abstain or stop.
The result may be a bounded pilot, a narrower connector - or a decision to fix source permissions before adding AI search.
This is exactly the scope of our n8n AI Automation service.
Frequently asked questions
It is the starting point. The retrieval layer must carry identity and control metadata, keep them current, authorise sensitive candidates and prevent unauthorised text from reaching the model.
They help narrow candidates, but mirrored ACLs can become stale. Sensitive domains may require a current source-backed authorisation check before content enters model context.
Only if its broad access and the downstream authorisation model are deliberately designed and tested. One powerful account can silently flatten boundaries between departments and clients.
Fail closed for the affected scope: abstain, expose the control failure and repair or reconcile the permission state. Do not treat missing permission evidence as permission granted.
Sources and evidence boundary
Research checked on 28 July 2026 against the primary or official sources linked in the article. The permission-chain table, metadata schema and Atlas scenario are IZZY guidance.
No tenant, identity provider, source ACL, connector token, vector store, prompt, cache, deletion path, retention schedule, n8n workflow, access denial or business outcome was tested. This article is technical and operational guidance, not a compliance or legal opinion. Product documentation can change after the research date.