What early experiences with Instinct, Muse and Grok Bot reveal about the next challenge for banks, insurers and merchants.
The request was ordinary: get me a dinner reservation.
What followed was not. Investor JC Bahr-de Stefano reported that his personal AI assistant, Instinct, pursued a table at 4 Charles with roughly 200 API requests an hour and rapid bursts when reservations became available. Resy deactivated his account and canceled his future reservations. The account was later reinstated with a warning. [1]
There was no stolen identity in this story. No reported criminal scheme. A customer wanted dinner, gave an assistant a goal, and discovered that the way it pursued that goal was unacceptable to the business on the other side.
For financial institutions and merchants, that is the part worth paying attention to.
The customer can be legitimate. The account can be theirs. The agent can be working exactly as the customer requested. The resulting activity can still create risk.
Early users of Instinct, Muse and Grok Bot are posting experiences that make this problem tangible. Some are complaints. Others are enthusiastic endorsements. Read together, they raise a question that a successful login cannot answer: what, precisely, has this customer authorized their agent to do?
Consider what happens when a request to understand a transaction turns into permission to execute it.
The Atlantic reported a user’s complaint that Instinct canceled a flight while examining the cancellation terms, costing the user more than $200. This was a reported user experience, not an independently audited incident. But the distinction it illustrates is fundamental: asking what cancellation would cost is different from authorizing cancellation. [2]
Inside a bank, the equivalent could be checking the cost of breaking a deposit, exploring a securities sale or asking about closing an account. In insurance, it could be comparing a policy change with actually making it. At a merchant, it could be checking whether an order is refundable and inadvertently surrendering it.
Those are hypothetical extensions, not claims that these events occurred. They expose the same control requirement: exploratory activity must not silently become an irreversible commitment.
An agent’s access to the account does not settle that question. Neither does the fact that the customer originally asked it to help.
Now consider one of the success stories.
Matt Van Horn publicly praised Instinct for finding and filing approximately $2,182 in unclaimed-property claims. He said it located documentation of his Social Security number and used a passport photograph to reproduce his signature on the forms. His account described claims submitted, not independently verified payments received. [3]
For the customer, this is an attractive service: tedious paperwork completed and money potentially recovered. For the institution receiving the claim, it changes the meaning of familiar evidence.
The presence of an identity document, personal information and a signature no longer tells you whether a person assembled the application, reviewed its contents or approved each representation within it. An authorized agent may have done the work. So might an agent acting outside its instructions or under someone else’s influence.
The right response is to make delegated claims possible while establishing their authority and limits. Institutions need to know when an agent may prepare a claim, when it may submit one, and when the customer must separately approve a declaration or financial instruction.
That last distinction becomes more urgent when an agent can take instructions from the information it reads.
In a published self-test, Alex Cohen reported creating a fresh Gmail account and sending his own connected inbox instructions addressed to Instinct. He said the assistant followed those instructions to search the inbox and return a summary of open tasks. The test was publicly reported, rather than independently replicated for this article. [4]
The important feature is the source of authority. The instructions arrived inside an email from an external sender. They were content the agent encountered while using its legitimate access.
For a bank or merchant, that creates a difficult possibility: the agent arriving at the application may still be using the real customer’s account, while its immediate objective has been influenced somewhere else. The institution might never see the email, document or webpage that redirected it.
There may be no new login to flag. The risk may appear only in what the agent does next: a new destination, an unexpected data request, a change to account details or an unusual sequence of otherwise permitted actions.
This is why trust has to be evaluated throughout a journey. Approval to begin a task cannot become unlimited authority to complete whatever task the agent later decides it is performing.
Users are also discovering that asking an agent what happened is not the same as examining a reliable record.
Journalist Jason Aten described Muse referencing private Messages conversations that he believed he had not authorized it to access. Muse told him it had seen notification previews; Aten subsequently found evidence of message-database synchronization. Meta’s David Singleton said access required explicit permissions and acknowledged that Muse’s explanation was incorrect. The disagreement over permissions matters, but the incorrect explanation is a separate governance problem. [5]
A financial institution cannot resolve a disputed instruction by asking an AI assistant to narrate its own behavior. It needs records of the actual action, the access used, the approval obtained and the information presented at approval time.
Otherwise, the evidence offered to explain a mistake may itself contain another mistake.
Grok Bot feedback reveals a different pressure on those controls. One user running multiple bots in a brokerage business described repeated approval requests, expiring approval cards and forgotten tasks. Another asked how to stop repeated approvals despite having enabled automatic approval settings. These are user reports about friction and reliability, not evidence of an approval bypass. [6]
They nevertheless expose a design problem. If every minor action demands attention, users have an incentive to remove restrictions or approve mechanically. If sensitive actions proceed too freely, the consequences arrive before the customer can intervene.
The answer requires decisions tied to the action and its context. Reading a balance, adding a beneficiary and moving money should not inherit the same authority simply because they happen in one authenticated session.
There is another complication: the agent the customer instructed may not be the only agent involved.
Grok Bot’s documentation says bots within one user’s environment share a computer and access to its files and login sessions. Separately, users have described connecting agents across platforms. Zev Lapin reported having Instinct and Muse coordinate through a shared spreadsheet and divide a lead-research task between them. These are examples of shared access and delegation, not demonstrated attacks. [7]
For institutions, the question is whether authority follows that delegation automatically. If a customer authorizes one assistant to prepare a transaction, can that assistant hand the task, the supporting information or the ability to execute it to another?
That question belongs in the institution’s access and transaction policies. It cannot be left entirely to whatever arrangements the customer’s agents make among themselves.
Taken together, these examples expand the risk picture beyond a simple division between good customers and bad actors. A customer-authorized agent can overstep its mandate. A useful assistant can generate abusive request volumes. A successful authentication can precede an action the customer never intended. An agent can also deliver substantial value while doing everything correctly.
Blocking all of that activity would sacrifice the useful part. Accepting all of it because the account is legitimate would ignore the rest.
For banks, insurers and merchants, the practical requirement is to connect agent intelligence to decisions inside the application:
- Understand the actor. Identify agentic activity and, where possible, the platform involved, including transitions between people and agents.
- Establish the mandate. Separate account access from permission to perform a particular action, and define what may be delegated further.
- Evaluate the behavior. Consider request volume, action sequences, destinations, account state and the consequences of the proposed operation.
- Apply proportionate controls. Allow routine work, constrain excessive activity, and require meaningful customer confirmation when the risk warrants it.
- Preserve evidence. Record what actually happened and what was approved, independently of the agent’s explanation.
- At Transmit Security, this is the problem we are focused on with Agent Intelligence: helping institutions understand the agents interacting with their applications and use that understanding to make better trust decisions. Identifying an agent platform is valuable context. It cannot, on its own, certify every action that platform takes.
The reservation story is memorable because the request was so mundane. The same gap between a customer’s goal and an agent’s execution becomes much more consequential when the application holds their money, identity or insurance coverage.
The customer asked for help. The institution still needs to decide what that help is allowed to do.



