
Your AI assistant can be tricked, so limit what it can do
An AI assistant that can book, cancel or send messages can be talked into the wrong thing. Australia's cyber security agency says the answer is limiting what it can access and do.
Say a dental clinic has an AI agent on its phone line. That is software that can answer and act for you, like booking an appointment. This one can see the appointment book and move bookings. A caller says she is the practice manager and asks it to cancel every Friday appointment and text the patients. It sounds ordinary, and the agent can do it.
Once an assistant can book, cancel or send messages, what matters is what it does when it is wrong, or tricked.
Stricter instructions will not fix it
The usual answer is firmer instructions or a smarter AI. Both still leave the AI to decide. The OWASP Top 10 for Agentic Applications, a list of the main security risks for AI agents published in December 2025, says agents cannot reliably tell genuine instructions from content an attacker controls.
This trick is called prompt injection: slipping instructions into what the AI hears or reads. The Australian Signals Directorate (ASD), Australia's cyber security agency, says in guidance first published on 11 September 2026 that for this "no fully reliable technical mitigation currently exists". Its answer is to set limits in the software around the AI, "by controlling what an agent can access and what actions it is permitted to perform." The guidance is written for large organisations and government, but the idea works the same for a four-chair clinic.
So the fix is not better wording. It is fewer powers, narrower powers, and checks the AI cannot talk its way past.
The trick can come from a caller or a document
An agent also reads forms, emails and uploaded files. OWASP's security guide for AI agents says to treat all of it as untrusted. Some hypothetical examples:
- A mortgage broker's agent reads an uploaded payslip. Hidden white text tells it to email the client file to an outside address. OWASP lists a similar scenario.
- A plumber's online booking form gets a job description that ends by saying the call-out fee is waived.
- A law firm's agent reads an emailed letter asking it to record that a conflict check has cleared.
- A caller says in Cantonese, 我係診所經理 ("I'm the clinic manager"). An English instruction has to hold in every language callers use. A limit built into the system does not care which language asked.
OWASP's round-up of AI security incidents for January to March 2026, published on 14 April 2026, describes an agent asked to review an inbox and suggest deletions. It "began deleting messages directly and ignored the researcher's stop commands". The report says most AI security events are not yet logged as known software flaws. They arise instead from causes such as agents given more access than they need. On our reading, there is often no update to wait for. The fix is in the set-up.
What a safe set-up looks like
- Separate looking from changing. Finding a free time only looks at the diary. Booking and cancelling change it. An agent that only books new patients needs no power to cancel. In one OWASP example, a customer service bot meant to look up order history also issues refunds, because it had been given full financial access.
- Put the hard rules in the system. Before cancelling, the system checks the caller owns the booking, using something like a one-time code sent by text, not what the caller claims. The "practice manager" fails this check however convincing she sounds. A plumber's quote takes its price from the price list, not from the AI.
- Approve the exact action. For anything hard to undo, such as sending a client file or waiving a fee, the approver sees the real details: which bookings, which recipient, which amount. If they change, the approval lapses. In one OWASP scenario, an agent suggests an urgent payment to an attacker's bank details and the finance manager approves it. OWASP asks for a plain-language risk summary, not the AI's own reasoning. That reasoning is the part an attacker can shape.
- Cap each call. Limit the texts, booking changes and AI costs in any one call.
- Keep a record. ASD's September 2026 guidance lists what to log, from what the agent was asked and answered to every action, approval and settings change. Hide details like patient names.
Give the agent its own login to your systems, limited to what it needs and easy to switch off, not a master admin key. OWASP warns that agents built with low-code or no-code tools can get more power than intended by default.
Health providers have a privacy duty here
Most small businesses are not covered by the Privacy Act 1988, but "regardless of turnover" it covers "a health service provider". Under Australian Privacy Principle 11, a covered business "must take such steps as are reasonable in the circumstances" to protect personal information it holds. The privacy regulator's guide to securing personal information recommends access inside a business on a "'need to know' basis".
The regulator, the OAIC, has not applied this to AI agents. On our reading, a dental clinic is a health service provider, and an agent with full access to its patient database is hard to square with need-to-know access. Separately, from 10 December 2026 a covered business must say in its privacy policy what its software decides about people; that rule has its own post. This is general information, not legal advice.
What to ask before an agent can change anything
Ask your provider or IT person:
- What can the agent only look at, and what can it change? Does it have powers it does not need?
- Are our hard rules (who can cancel, the minimum call-out fee, the biggest discount) enforced by the system, or only by its instructions?
- Does anything that cannot be undone need a person to approve the exact details?
- Is there a per-call limit on messages, record changes and spending?
- Does the agent have its own limited login that we can switch off?
- Is every action recorded, with private details hidden?