Treat public clues as leads, not findings
An external review can show what an organisation makes visible to an unauthenticated observer. It cannot, by itself, establish that a system is deployed, reachable, vulnerable, or compromised. A provider name in a browser policy might be unused. A certificate name might belong to a retired service. A job description might describe a planned architecture. Keep those distinctions in the ticket from the first observation.
This checklist closes the AI Risk Stories series with five passive sources: Content-Security-Policy (CSP), certificate transparency (CT) logs, job postings, public repositories, and robots.txt. The first four can disclose useful architecture clues. robots.txt provides supporting context about crawler preferences; it is not an inventory or a security boundary.
A public signal does not prove deployment
Record exactly what you observed, when you observed it, and where it came from. Do not label a signal as active use, exposure, or compromise until an authorised owner confirms it with current internal evidence.
Start with a bounded, authorised review
Define the domains, repositories, and public recruiting sources in scope before collecting anything. Get written approval from the asset owner, follow the organisation’s rules for third-party properties, and use ordinary public pages and records only. This is a review of public material, not permission to probe systems.
- Set a time window and record the reviewer, approved scope, and source URLs.
- Use low-impact requests to public pages or public records. Do not authenticate, enumerate hidden paths, bypass controls, submit forms, test credentials, or send traffic to discovered endpoints.
- Save the minimum evidence needed to reproduce the observation, such as a response header excerpt, public certificate record, job URL, or repository path. Redact tokens, personal data, and unrelated content.
- Do not copy secrets into tickets or chat. If a public repository appears to contain a credential, restrict access to the finding, notify the repository owner through the approved incident channel, and follow the secret-response process.
- Stop and hand off if the review reveals suspected live credentials, personal data, or a service that appears unintentionally exposed.
Use the five signals to form testable questions
Each source has a different failure mode. Capture the raw observation separately from your interpretation, then ask an internal owner to verify whether it describes a current system.
- CSP: A provider domain in connect-src means browsers are permitted to connect to that origin under the policy. It does not prove the page makes a call, that a user has a credential, or that the organisation sends data directly. Ask the web platform owner to compare the header with current frontend code and network architecture.
- Certificate transparency: A certificate entry records a name covered by a publicly logged certificate. It can reveal a subdomain naming pattern, but not whether DNS resolves now, whether a service responds, or whether it is an AI endpoint. Ask DNS and platform owners to reconcile candidates with current asset inventory and retirement records.
- Job postings: A role description can name model providers, frameworks, cloud services, or data types. Recruiting copy can be aspirational, stale, or broad. Ask the hiring manager or engineering owner which items are approved and present in production, staging, or neither.
- Public repositories: A repository can expose architecture, dependency choices, sample configuration, or accidentally committed secrets. A public code reference alone does not prove the code is built or deployed. Ask the repository owner to check commit history and deployment references; route possible secrets through the restricted incident process.
- robots.txt: A rule documents crawler preferences for a user-agent and path. It does not block access, authenticate a requester, or prove that a named path contains an AI system. Ask the web owner whether the rules are intentional and whether sensitive paths have real access controls.
Keep observation and interpretation separate
Write “CSP permits connections to api.example” as the observation. Write “the frontend calls this provider” only after code or traffic evidence confirms that behaviour in an authorised review.
Triage by evidence strength and consequence
Assign one confirmation level to each item. Raise priority for evidence of exposed sensitive data or a live credential, not because a signal looks unusual. A low-confidence lead can still deserve a timely owner check when the potential consequence is high.
| Signal and initial level | What would confirm it | Owner | Safe next action |
|---|---|---|---|
| CSP provider domain Lead only | Current frontend code or approved browser testing shows a real request and identifies the data path. | Web platform or application security | Review the integration and data classification. Confirm whether a server-side gateway or approved browser call is intended. |
| CT name resembling an agent or model service Lead only | Asset inventory and service owner confirm the hostname and its current environment and purpose. | DNS, cloud platform, or service owner | Reconcile the name with owned assets and decommissioning records. Do not probe the host from an unauthorised network. |
| AI stack named in a job posting Lead only | Hiring manager or technical owner confirms an approved or active project and its environment. | Hiring manager and AI platform owner | Check whether the posting is current and whether its technical detail exceeds recruiting guidance. Correct only what the owner verifies. |
| AI integration or credential-like value in a public repository Potential exposure | Repository owner validates the finding in a restricted channel; for a possible secret, the credential owner checks validity without disclosing it. | Repository owner and incident response; credential owner for secrets | Limit circulation, follow incident response, and revoke or rotate a confirmed credential. Removing a file alone does not invalidate a secret. |
| robots.txt names an AI-related path or crawler Context only | Web owner confirms the path’s purpose and separately verifies its access controls. | Web platform or content operations | Treat disallow rules as crawler guidance. Protect private content with authentication and server-side authorization. |
Apply these confirmation levels consistently: Lead means a public clue exists but internal relevance is unverified. Corroborated means an asset owner ties it to a current system or approved project. Confirmed exposure means an authorised review verifies an unintended access path, sensitive data disclosure, or valid credential. Incident means the response team has accepted a security event under its incident criteria. A signal should not skip levels just because its wording sounds definitive.
Turn the review into owned remediation
Create one finding per distinct issue. Include the source, timestamp, exact observation, confidence level, business consequence under review, owner, due date, and the evidence needed to close it. Avoid speculative titles such as “AI endpoint exposed” when the evidence only shows a certificate name.
- For a confirmed browser-to-provider integration, document what data the client sends, which identity controls access, how secrets are kept out of browser code, and which team approves the integration. Remove unused CSP entries after the application owner verifies they serve no required traffic.
- For an unowned CT name, ask DNS and platform owners to identify the service before changing records. If it is retired, follow the normal decommissioning path and verify the associated DNS, certificate renewal, and ownership records.
- For a verified repository secret, revoke or rotate it first, inspect authorised usage logs as incident response directs, and then clean the repository history and downstream copies. Treat deletion from the latest commit as insufficient remediation.
- For confirmed recruiting overshare, have the technical owner and recruiting team remove implementation details that are not needed for the role. Do not ask recruiting to conceal ordinary technology use as a substitute for securing the system.
- For a path listed in robots.txt, verify server-side authorization independently. If content is meant to be public, reconsider whether the rule is useful; if it is private, enforce access control regardless of crawler behaviour.
Close with internal evidence
Close a lead when the owner records why it is stale, unrelated, or expected. Close a confirmed issue when the owner supplies evidence that the specific control changed and the risk review accepts the result. Keep the original observation attached so the next review can distinguish recurrence from a new signal.
Make the next review easier to defend
A useful AI risk surface review is a traceable chain from public observation to an owner’s decision. It does not turn an unfamiliar hostname into an incident, and it does not treat an absence of public clues as proof that no AI system exists. Keep scope, confidence, and internal confirmation visible in every finding. That record gives security and engineering a concrete basis for deciding what to fix and what to document as expected.
