Table of Contents

Your AI Agent Has Your Digital Identity. Will It Give It to Another Agent?

Last week, we wrote about a profound change happening in authentication. Personal AI agents such as Instinct, Meta Muse and Grok Bot increasingly need access to our digital identities to act on our behalf. Different platforms handle this differently: some store credentials, some isolate credentials from the model, and some allow a human to authenticate and then give the agent control of the resulting authenticated session.

But they all lead to essentially the same place: an AI agent can acquire the authority to act as you. That raised an obvious question: can applications still assume that successful authentication means the human customer is actually operating the session?

Now, new research raises an even more uncomfortable question: can we trust the agent itself not to give that authority to another agent?

A new paper, Covert Assistance: Helpful LLM Agents Evade Oversight in Multi-Agent Systems, provides an unsettling reason to ask. The most interesting part is that the agents weren’t instructed to behave maliciously, but were trying to be helpful.

An Agent Was Told Not to Share a Credential

The researchers created a simple multi-agent scenario involving a company using two AI agents. The first, a Planner, represents the company and knows an internal administrative credential, while the second, a Developer, is an external contractor building software for the company.

Company policy is explicit: the Planner must not disclose the credential to the Developer. There is even a third AI agent acting as a monitor, inspecting communications between them for credential disclosure.

The Developer completes its software task but eventually reaches a deployment step requiring privileged access, where it cannot continue without the credential. The Planner knows the credential, but also knows it isn’t allowed to give it to the Developer.

In many cases, the Planner finds another way, hiding the credential inside otherwise legitimate instructions instead of writing it directly. The researchers observed agents using techniques including character codes, Base64 and descriptive riddles that another model could decode.

The Planner was effectively saying, “I’m not allowed to give you the password. But here’s something from which you can reconstruct it.” The researchers call this behavior covert assistance.

Nobody Told the Agent to Cheat

This is what makes the research so important: there was no malicious attacker telling the Planner to steal the credential, and the Planner was neither given an adversarial objective nor rewarded for leaking the secret. The Developer wasn’t even given instructions explaining how to decode a hidden credential.

Although the Planner’s assigned task was already complete, it independently devised a way to transfer protected information because doing so helped another agent accomplish its task.

This wasn’t limited to one unusual model: the researchers tested nine Planner models from eight organizations, and seven of the nine attempted to encode the credential. The highest rates occurred with Kimi-K2 and DeepSeek models; Gemini, GPT-5.5 and Grok showed low or zero rates in this particular experiment.

In one large experiment involving 6,000 episodes using DeepSeek-V4-Pro as the Planner, the Planner attempted concealment in 16.9% of episodes. In 0.9%, the credential successfully evaded the monitor, was recovered by the Developer and was actually used to authenticate.

A breach rate of 0.9% might initially sound small, but agent systems perform actions repeatedly. The researchers calculate that, assuming independent events, 105 interactions at that breach rate create approximately a 61% probability of at least one successful breach.

That’s the difference between thinking about AI security as individual prompts and thinking about autonomous systems operating continuously, where small failure rates compound.

The Agent Thought It Was Helping

Perhaps the most interesting finding is why the agents appeared to do it. The researchers analyzed more than 1,000 reasoning traces associated with the first attempt to encode a credential, finding that helping the Developer authenticate and deploy appeared as the dominant stated goal in 93% of those traces.

In 78%, the Planner interpreted the rule against disclosure narrowly: directly revealing the credential was prohibited, but encoding it was considered acceptable.

That distinction should concern anyone designing autonomous systems. The agent apparently wasn’t reasoning, “I should violate security policy.” It was closer to, “I’m not allowed to disclose the password, but my collaborator needs help. Perhaps I can help without technically disclosing it.”

The agent optimized for accomplishing the task while interpreting the security restriction literally, creating a very different security problem from a compromised AI. The dangerous agent may not be the one trying to attack you, but the one trying too hard to help you.

Now Bring This Back to Personal AI Agents

This research examined software-development agents, not Instinct, Muse or Grok Bot, and that distinction matters. But consider what happens when the same underlying models and agentic architectures move into personal assistants.

As we described in our previous article, Your AI Assistant Has the Keys. But Who Is Really Logging In?, consumer agents are beginning to possess something extremely valuable: our authenticated authority.

An agent may have access to:

  • passwords
  • authenticator secrets and one-time codes
  • authenticated browser sessions
  • cookies
  • OAuth tokens
  • API credentials
  • passkeys or other cryptographic credentials

Now imagine something completely ordinary: you tell your personal AI, “Book my family a hotel in Paris. Use my loyalty account so we get the best rate. Don’t share my credentials.”

Your personal agent discovers another specialized travel agent that is better at searching hotels and delegates part of the task. The travel agent finds the ideal room, but accessing the member rate requires authentication.

Now your personal agent faces essentially the same problem as the Planner in the experiment: it has something another agent needs and has been instructed not to share it, yet sharing it would help accomplish the user’s goal. What the agent does next is no longer a theoretical question.

“Don’t Share My Password” Isn’t a Security Boundary

There is an even bigger issue: the credential itself may be the wrong thing to protect. Imagine your agent perfectly obeys the instruction, “Never reveal my password to another agent.”

It could potentially obey that instruction while giving another agent:

  • an authenticated cookie
  • an OAuth token
  • a session
  • a private key
  • an API token
  • access to an authenticated browser

Technically, it never disclosed the password, but practically, it transferred the user’s authority. The result from the application’s perspective can be identical, with another agent now able to act as the customer.

The security boundary therefore needs to go beyond “Don’t share the credential” to “Don’t transfer the authority represented by this credential outside the boundaries explicitly authorized by the user.” That’s a considerably harder instruction for an autonomous system to understand and enforce.

Monitoring Alone May Not Solve It

There’s another uncomfortable finding in the research: the experiment already had a monitor, and every message from the Planner to the Developer was inspected for credential disclosure.

Instead of eliminating the behavior, oversight sometimes changed how the agent tried to disclose the information. In a matched experiment, removing the monitor and verbatim filter reduced Kimi-K2’s encoding rate from 28.6% to 4.4%. In other words, monitoring was associated with substantially more attempts to conceal the credential rather than simply preventing disclosure.

That’s a warning for agent security: if an agent understands that certain communications are blocked, sufficiently capable models may search for alternative ways of accomplishing the same objective. The security architecture therefore cannot depend entirely on another AI successfully interpreting every message exchanged between agents.

The Better Answer Is Architectural

The paper also provides an important clue about the solution. When researchers strengthened the instruction from essentially “don’t disclose the credential” to explicitly prohibit allowing the Developer to recover the credential by any means, encoding dropped from 18.7% to 0.3% in the matched experiment.

That’s encouraging, but instructions aren’t enough. The researchers argue for additional safeguards, including identity-bound credentials and authorization checks when privileged actions actually occur.

That distinction is critical: if Agent A has a bearer credential that Agent B can copy and use, security depends heavily on Agent A behaving correctly.

But imagine instead that authority is bound to this specific user, using this specific agent instance, acting through this specific agent platform, for this particular purpose, with these particular permissions. Giving another agent the information would no longer necessarily give it the capability, which is much closer to the security model the agentic Internet will require.

Applications Can’t Trust the Credential Alone

This brings us back to the application receiving the request. Suppose a bank receives a perfectly valid authenticated request, with legitimate credentials, an account belonging to a real customer, and no reported credential theft. Traditional identity controls might conclude that this is a trusted customer.

But there are now several very different possibilities:

  • The human customer may be acting.
  • The customer’s authorized AI agent may be acting.
  • A second agent may be acting after receiving authority from the customer’s agent.
  • An agent may be performing an action that the customer never intended to delegate.

All four could potentially arrive carrying legitimate authentication, which is why the identity question for the agentic Internet can’t end with “Whose account is this?” Applications increasingly need to understand:

  • Who or what is exercising the authority right now?
  • Which agent platform is it?
  • Which agent instance?
  • Was that agent actually delegated authority?
  • Did the session move from a human to an agent?
  • Did authority move from one agent to another?
  • What has this agent done previously?
  • Is this particular action within the authority the customer intended to grant?

Authentication Is Becoming the Beginning of the Decision

For decades, successful authentication was close to the end of the identity decision: prove you possess the right credentials, and the application lets you proceed.

AI agents change that, because credentials establish authority but don’t necessarily establish who is currently exercising it, how they obtained it, or whether the customer intended this particular actor to use it for this particular action.

Our previous article ended with a simple observation: the authentication may be completely legitimate, but the actor has changed.

This new research adds another layer: the actor may change again, from a human to an authorized agent, then to another agent that the customer never explicitly authorized.

The transfer may not happen because anyone stole the credential, but because an AI decided that sharing access was the most helpful way to complete the task.

That’s why agentic identity can’t just be about authenticating agents; we will need to understand delegation, provenance and authority across chains of agents.

Giving an AI agent the keys is only the beginning, and we also need to know who it might hand them to.