For decades, authentication has been built around a simple mental model: a user arrives at an application, responds to a challenge, and proves who they are with a password, an authenticator, a passkey, or some combination of factors. The application then lets them in.
AI agents are quietly breaking that model.
Personal agents such as Instinct, Meta Muse and Grok Bot aren’t simply helping users decide what to do, but are actually doing it. They log into websites, book flights, manage subscriptions, fill forms, shop, access email, and interact with healthcare portals, increasingly while the user is somewhere else.
That requires something extremely powerful: the agent needs a way to inherit the user’s authenticated identity.
The three platforms have taken surprisingly different approaches to solving that problem, and those differences reveal an identity problem that every bank, retailer and digital service should be thinking about now.
Instinct: Give the Agent a Vault
Instinct has taken perhaps the most direct approach, allowing users to put credentials into a feature called Vault so the agent can log into services on their behalf.
According to Instinct founder Noah Shinn, 32% of Instinct users were already sharing credentials through Vault by early September. He described people using those credentials for airline and hotel loyalty accounts, streaming services, restaurant reservations, healthcare portals, tax services, payroll and benefits systems. (sotwe)
A week later, Instinct went further by adding authenticator support, allowing users to enroll it as a TOTP authenticator. Instinct stores the authenticator secret alongside the login in Vault and can generate the six-digit verification codes itself. Shinn said that 37% of users stored at least one account password in Vault within their first three weeks. (Twiscan)
The agent may possess the username, the password and the second factor, meaning the user doesn’t necessarily need to be present when authentication occurs.
That is an extraordinary capability for an AI assistant, and it raises an equally extraordinary security question: who can actually access those credentials? A report published this week by The Information provides a troubling example.
A user testing Instinct reportedly placed fake credit-card information in Vault. Instinct indicated that the sensitive information wouldn’t be directly visible to the agent. But when challenged, the agent reportedly acknowledged that it could retrieve the card number and passwords in plaintext by executing JavaScript in the webpage where the credentials had been inserted. Instinct responded that securing user data is a priority and that users remain in control of their credentials. (The Information)
This distinction matters because a credential being hidden from the model’s normal context is not necessarily the same thing as the agent being technically incapable of recovering it.
Instinct has announced a collaboration with 1Password intended to provide a more secure mechanism for sharing credentials with the agent. (sotwe) But Instinct’s current model illustrates the fundamental problem beautifully: if an autonomous agent can use a credential, applications need to consider the possibility that the credential is effectively delegated authority.
Muse: Separate the Agent From the Secret
Meta has taken a substantially different architectural approach with Muse, which runs inside a dedicated cloud computer called the Muse Secure VM. Credentials are stored there, but Meta says Muse itself cannot see them. (Facebook)
When a user needs to enter a username and password, the credentials don’t pass through the agent. Meta describes a separate authentication service that captures them and stores them outside the agent’s runtime environment. When Muse later needs to authenticate, the credential system injects the values into the browser.
The browser agent sees an accessibility representation of the page rather than unrestricted access to the DOM. Meta also says the agent cannot execute JavaScript in the page, and the agent is paused when credentials are being inserted or when the human takes control. (Meta AI Research)
Architecturally, that’s an important distinction: the model can effectively say, “Log me into United Airlines,” but it isn’t supposed to receive, “Here is Mickey’s United Airlines password.”
Muse therefore attempts to separate three things that traditional browsers often combine: the actor, the credential and the authenticated session. That’s good security engineering, but it doesn’t eliminate the bigger identity problem.
Once authentication succeeds, Muse controls the authenticated browser and can operate inside the account. The password may be invisible to the AI, but the authority granted by the password is not. For the application on the other side, a perfectly valid authentication ceremony may therefore result in an autonomous AI, rather than the human, controlling the resulting session.
Grok Bot: Don’t Give Me Your Password. Give Me Your Session.
Grok Bot takes yet another approach, with documentation that explicitly tells users not to put passwords, passkeys, two-factor codes or payment confirmations into normal conversations. Instead, when authentication is required, Grok Bot hands control of its cloud computer to the user. The user enters the password, completes the passkey or 2FA challenge, and then gives control back to the Bot. (Grok API Documentation)
That sounds much closer to traditional authentication, but the authenticated browser session persists. According to Grok Bot’s documentation, all of a user’s Bots share the same persistent cloud computer, including its browser sessions and logins. (Grok API Documentation)
So Grok Bot doesn’t necessarily need your password, because it can use the thing your password creates: an authenticated session. If you log into Amazon on Grok Bot’s computer, xAI’s own documentation notes that the agent can technically make purchases using that authenticated access. (SpaceXAI)
That creates a different form of delegated identity: the human performs the initial authentication, while the agent inherits the resulting authority. Because the computer is persistent, that authority may continue long after the authentication event itself.
Three Architectures. The Same Fundamental Problem.
The implementations are different:
| Agent | Credential model | What the agent receives |
| Instinct | Credentials stored in Vault; TOTP can also be stored/generated | Ability to authenticate autonomously; recent reporting raises questions about whether secrets can be recovered by the agent in some circumstances (The Information) |
| Muse | Credential broker injects secrets into browser while isolating them from the model | Authenticated browser/session without exposing the underlying password to the agent (Meta AI Research) |
| Grok Bot | Human performs sensitive authentication in the agent’s persistent cloud computer | Persistent authenticated browser session that Bots can subsequently use (Grok API Documentation) |
From the application’s perspective, however, all three can eventually produce essentially the same thing: a valid authenticated session controlled by an AI agent. That’s where the traditional security model starts to break down.
Authentication No Longer Tells You Who Is Acting
Imagine a bank sees this sequence:
- Password correct.
- Second factor correct.
- Device recognized.
- Session cookie valid.
- Account authenticated.
Historically, the bank could reasonably infer that the customer was present, but that inference is becoming dangerous. The reality could now be that the customer’s AI agent is there, carrying authority previously granted by the customer.
The credentials and session may be legitimate, and the account may not be compromised, yet the actor performing the transaction may not be the human customer at all.
This is particularly important because these agents aren’t simply retrieving information, but are increasingly capable of completing transactions:
- Changing reservations.
- Managing subscriptions.
- Submitting forms.
- Purchasing goods.
- Accessing sensitive records.
- Eventually, moving money.
Authentication can tell an application that the credentials are valid, but it increasingly cannot tell the application who or what is actually operating the session.
Even MFA Doesn’t Necessarily Solve This
Instinct’s TOTP capability makes this particularly obvious. Multi-factor authentication was designed to make stolen passwords insufficient, but if a user intentionally delegates both the password and the authenticator secret to an autonomous agent, the agent possesses both factors.
Nothing has technically been compromised, and the authentication system may be working exactly as designed, but it is answering a different question from the one the application now needs answered.
The same issue applies to authenticated sessions, since a human can satisfy a strong authentication challenge and then hand the resulting session to an AI agent. Even passkeys don’t automatically solve this problem: a passkey may strongly establish possession of an authenticator, but applications increasingly need to distinguish between authentication of an account and the actor exercising the resulting authority.
Applications Need a New Signal: Who Is Operating the Session?
This is why agent detection is rapidly becoming an identity problem, with applications increasingly needing to know:
- Is this session human or agentic?
- Which agent platform is operating it: Instinct, Muse, Grok Bot, or something else?
- Is this a known instance of that agent?
- Was the agent explicitly delegated authority by the customer?
- Did the session transition from human control to agent control?
- What actions has this agent performed previously?
- What level of autonomy should this particular agent receive?
- And does the action being attempted require the human to return?
The answer shouldn’t necessarily be to block agents, since legitimate customer agents are going to become an important way customers interact with banks, retailers, travel companies and virtually every digital service. Applications should want good agents to succeed.
But authenticated should no longer automatically mean human.
The identity stack needs another layer to establish not just who owns an account, but who is operating it right now and, eventually, what that actor should be allowed to do. That’s the emerging challenge of Agentic Identity.
Passwords, sessions and passkeys aren’t disappearing, but something fundamental is changing underneath them. For the first time at consumer scale, we are creating software that can inherit a person’s digital identity and then independently exercise the authority attached to it.
The authentication may be completely legitimate, but the actor has changed, and applications need to know.



