Business Strategy

Azure AI Services Pricing in Australia: 2026 Cost Guide

Azure AI Services Pricing in Australia: 2026 Cost Guide

Abstract illustration of cloud AI services connected to a budget gauge and stacked coins over a faint map of Australia

Model tokens are rarely the biggest line on the invoice

Most Azure AI budgets in Australia start with the wrong number. Someone opens the Azure OpenAI price list, sees a mini model at about a dollar per million input tokens, multiplies by an optimistic usage guess and gets a figure small enough to approve on a credit card. In the hypothetical knowledge assistant worked through later in this post, model tokens come to about a fifth of the monthly bill. The search index, the second search replica you need for an uptime commitment, the log workspace capturing every prompt, and the non-production environment nobody switched off make up the rest.

This guide sets out Azure AI services pricing in Australia as it stood when we checked it in October 2026: Azure OpenAI in Microsoft Foundry, Azure AI Search, Document Intelligence, Speech, and the Microsoft 365 Copilot and Copilot Studio licences that often sit in the same budget conversation. It explains how each one bills, what Australian data residency does to the price, and works through two hypothetical monthly estimates line by line so you can swap in your own volumes.

How the prices in this post were checked. Every Azure unit price below comes from Microsoft's Azure Retail Prices API (the same data behind azure.microsoft.com/pricing and the Azure pricing calculator), queried for the Australia East region in AUD on 8 October 2026. Microsoft's AUD list prices on that date worked out at about A$1.4246 per US$1, and where Microsoft quotes only USD we convert at that same rate. Figures exclude GST and any enterprise agreement discount. Prices change often, so confirm current rates in the Azure pricing calculator before you commit a budget.

If you are still choosing between hyperscalers, start with our AWS vs Azure vs Google Cloud for AI in Australia comparison. This post assumes Azure is already the platform and the question now is what it will cost.


How an Azure AI bill is built

An Azure AI workload is several metered services wired together, each billing on a different unit. Knowing the unit is most of the battle, because it tells you which number to estimate.

Where the money goes in a typical Azure AI workload

Ingest
Blob storage per GB, Document Intelligence per 1,000 pages
Index
Embeddings per token, AI Search per hour per search unit
Generate
Model tokens per million, or PTUs per hour
Observe
Log Analytics per GB ingested and retained
Run
People time for governance, evaluation and support
ServiceWhat you pay forAustralia East list price (AUD, ex GST)
Azure OpenAI GPT-5.4 mini, Global StandardPer million tokensA$1.0685 input, A$0.1068 cached input, A$6.4107 output
Azure OpenAI GPT-4.1 mini, regional StandardPer million tokensA$0.63 input, A$0.16 cached input, A$2.51 output (converted from US$0.44, US$0.11, US$1.76)
text-embedding-3-small, regionalPer million tokensabout A$0.03 (converted from US$0.022)
Azure AI Search BasicPer search unit per hourA$0.1895 (about A$138 a month)
Azure AI Search Standard S1Per search unit per hourA$0.6325 (about A$462 a month)
Semantic ranker, standard planPer 1,000 queriesA$1.4246
Document Intelligence ReadPer 1,000 pagesA$2.1369 (A$0.8548 beyond 1 million pages)
Document Intelligence prebuilt modelsPer 1,000 pagesA$14.246
Document Intelligence custom extractionPer 1,000 pagesA$42.7381
Speech to text, real-time / batchPer audio hourA$1.4246 / A$0.2564
Neural text to speechPer million charactersA$21.369
Log Analytics ingestionPer GB after 5 GB free each monthA$4.7582
Blob storage, Hot LRSPer GB per monthA$0.0285
Internet data transfer outPer GB after the first 100 GB a monthA$0.171 (Microsoft network routing)

Source: Microsoft Azure Retail Prices API, Australia East, AUD, as checked October 2026. Monthly figures assume 730 hours.


Australian data residency changes the price, and the model list

This is the line item most budgets get wrong, and it sits in a dropdown labelled "deployment type".

Microsoft's deployment types documentation (updated August 2026) states that data stored at rest stays in your chosen Azure geography for every deployment type. Where the prompt is processed is a separate matter:

  • Global deployments may process prompts in any Azure region.
  • Data Zone deployments process within a Microsoft-defined zone. For this part of the world that is Asia Pacific, which covers several countries beyond Australia.
  • Standard (regional) and Regional Provisioned deployments process prompts within the Azure geography you choose.

The catch for Australian organisations is availability. On Microsoft's region availability tables (as checked October 2026), the pay-as-you-go regional Standard type in Australia East offers GPT-4.1 mini, GPT-4o and the embedding models. The GPT-5 family is available in Australia East on Global Standard, on Data Zone Standard, and on Regional Provisioned. Australia Southeast does not appear in those Azure OpenAI tables at all, which matters if you were planning a second Australian region for disaster recovery.

So the price of keeping prompt processing onshore depends on which model you need:

GPT-5.4 mini in Australia East: what each processing boundary costs

Metric
Global Standard (any region)
Onshore options
BillingPer token, A$1.0685 in / A$6.4107 out per 1MData Zone APAC: A$1.2821 in / A$7.6929 out per 1M (processing in APAC, not Australia only)
Processing inside AustraliaNot guaranteedRegional Provisioned only, minimum 25 PTUs
Cheapest onshore entry pointNo minimum spend25 PTUs x A$464.42 per PTU per month reservation = A$11,610.51 a month
Same deployment on hourly PTU billingNot applicable25 PTUs x A$3.2481 x 730 hours = A$59,277.82 a month
Onshore pay-per-token alternativeNot applicableOlder GPT-4.1 mini on regional Standard, A$0.63 in / A$2.51 out per 1M

Read that table carefully before a residency requirement gets written into a project charter. If your policy says "prompts must be processed in Australia" and you want a current-generation model, the floor is a provisioned reservation of roughly A$11,600 a month before you have answered a single question. If the policy is "data at rest stays in Australia and processing stays in Asia Pacific", Data Zone Standard costs about 20% more per token than Global and has no minimum. If you can work with GPT-4.1 mini, regional Standard keeps processing onshore on a per-token bill.

Which of those your organisation needs is a legal and risk question before it is a technical one. Our guide on where AI data actually goes covers the Privacy Act and contractual angles, and the private AI infrastructure page explains when hosting a model yourself becomes the cheaper way to meet a strict residency rule.

Choosing an Azure OpenAI deployment type

What does your data policy require for prompt processing?
No location restriction on processing, data at rest in Australia is enough
→ Global Standard: lowest token price, newest models first
Processing must stay in Asia Pacific
→ Data Zone Standard: about 20% more per token, no minimum
Processing must stay in Australia, an older mini model is acceptable
→ Regional Standard with GPT-4.1 mini, per-token billing
Processing must stay in Australia on a current model
→ Regional Provisioned: 25 PTU minimum, budget A$11.6K+ a month
Large overnight jobs with no latency need
→ Batch: half the Global Standard rate, 24-hour target turnaround

Pay-as-you-go or provisioned throughput

Provisioned throughput units (PTUs) reserve model capacity for you, billed per PTU per hour whether or not you send a request. Microsoft's provisioned throughput sizing guide (updated September 2026) lists, for GPT-5.4 mini, a Global and Data Zone minimum of 15 PTUs, a regional minimum of 25 PTUs, and 7,900 input tokens per minute per PTU, with one output token counting as six input tokens.

That lets you check break-even yourself. Fifteen Global PTUs on a monthly reservation cost 15 x A$370.40 = A$5,555.95 a month. At 100% utilisation around the clock for 30 days they process about 5.1 billion normalised tokens, which would cost roughly A$5,470 on Global Standard. In other words, a GPT-5.4 mini reservation only matches the pay-as-you-go price when it runs flat out every hour of the month. On hourly billing the same 15 PTUs cost A$15,599.37.

PTUs are worth paying for when you need predictable latency for a customer-facing or high-volume system, or when onshore processing on a current model is a hard requirement. For an internal assistant used during business hours, pay-as-you-go is nearly always cheaper. Microsoft also notes that a reservation does not guarantee capacity, so you create the deployment first and buy the reservation second.


The costs people miss

Prompt length, every time

Retrieval-augmented assistants send the system instructions, the retrieved document chunks and the conversation history with every question. A 6,000-token prompt for a one-line question is normal. Every chunk you stop retrieving comes off the input count of every question, so retrieval settings are a cost lever as well as a quality one.

Caching only works on identical openings

Per Microsoft's prompt caching documentation, a prompt must be at least 1,024 tokens long and its first 1,024 tokens must match exactly for a cache hit, and in-memory caches typically clear after 5 to 10 minutes of inactivity. Put stable instructions first and anything that varies (dates, user names) last. On GPT-5.6 and later models, cache writes can also be charged.

Reasoning tokens bill as output

GPT-5 family models generate internal reasoning tokens that are billed at the output rate. A higher reasoning effort setting can multiply output spend on the same question.

AI Search runs by the hour

A search service bills every hour it exists, used or not. Microsoft's service limits page says the uptime SLA for queries needs two or more replicas, which doubles the bill. The same page notes Basic does not support private connections to a Foundry resource for indexers, so a design that needs private networking for AI enrichment pushes you to S1 or higher.

Logging the conversation

Capturing full prompts and responses in Log Analytics for audit is sensible, and at A$4.76 per GB after the free 5 GB it adds up quickly on a chatty assistant. Decide the retention period on purpose.

Non-production environments

A dev and test copy of the stack carries its own search service and its own logs. A Basic search unit left running all month is A$138.

Data transfer

The first 100 GB a month of internet egress is free and AI responses are small, so egress rarely matters for a chat assistant. It starts to matter when you move large document sets between regions or out to another cloud.

People time

Someone has to own evaluation, prompt changes, access reviews, model version upgrades and the monthly cost review. That cost does not appear on the Azure invoice, and on a modest workload it can exceed everything that does. Our AI implementation cost breakdown covers the build and run labour in more detail.


Worked example 1: an internal knowledge assistant (hypothetical)

Consider a hypothetical organisation that puts its policies, procedures and contract templates behind an internal assistant. These assumptions are illustrative only:

  • 40,000 questions a month (for example, 100 regular users asking 20 questions on each of 20 working days)
  • 6,000 input tokens per question (instructions, retrieved chunks, history), a quarter of them served from cache
  • 500 output tokens per answer
  • 20,000 documents, about 50 million tokens to embed once
  • GPT-5.4 mini on Global Standard, AI Search S1 with two replicas for the SLA, semantic ranker on every query
  • 20 GB a month of logs, 50 GB of source documents, a Basic search unit for dev and test

Hypothetical knowledge assistant: monthly Azure cost (AUD, ex GST)

Model input, 180M uncached tokens x A$1.0685/MA$192.33
Model cached input, 60M tokens x A$0.1068/MA$6.41
Model output, 20M tokens x A$6.4107/MA$128.21
Embeddings, 52M tokens (one-off index plus queries)A$1.63
AI Search S1, 2 replicas x 730 hrs x A$0.6325A$923.45
Semantic ranker, 40,000 queries x A$1.4246/1,000A$56.98
Log Analytics, 15 GB billable x A$4.7582A$71.37
Blob storage, 50 GB x A$0.0285A$1.43
Dev and test search, Basic x 730 hrs x A$0.1895A$138.34
Estimated monthly Azure total (excludes app hosting and people)A$1,520.15

Model tokens are about 22% of this total. Search, at roughly 65%, is the line to optimise first. Two adjustments show how sensitive the estimate is:

  • Moving to AI Search Basic with two replicas brings search to A$276.67, if the index fits Basic's 15 GB storage and 5 GB vector quota per partition and you do not need private connections for enrichment.
  • Switching to GPT-4.1 mini on regional Standard keeps prompt processing in Australia and drops the token lines to about A$172.38, in exchange for an older and less capable model.

App hosting (App Service, Container Apps or a Copilot Studio front end) is left out because it varies so much by design.


Worked example 2: invoice and form processing (hypothetical)

Now a hypothetical finance team processing 10,000 supplier invoices a month, averaging two pages each, so 20,000 pages.

Three ways to extract 20,000 pages a month (hypothetical)

Metric
Approach
Monthly Azure cost and trade-off
Prebuilt invoice modelDocument Intelligence, 20 x A$14.246 per 1,000 pagesA$284.92. Least build effort, fields arrive structured
Prebuilt on a commitment tier20,000-page pre-built commitment tierA$270.67 a month flat. Only worth it if volume is steady
Read OCR plus a language modelRead at A$2.1369 per 1,000 pages, then GPT-5.4 mini on BatchA$42.74 + A$34.19 = A$76.93. Cheapest, but you build and test the extraction prompts
Custom extraction modelDocument Intelligence custom, 20 x A$42.7381 per 1,000 pagesA$854.76. For unusual layouts the prebuilt model misses

The language model line assumes 4,000 input and 400 output tokens per invoice on Global Batch (A$0.5342 and A$3.2054 per million). Batch suits this job because invoices can wait overnight, and Microsoft prices Batch at half the Global Standard rate with a 24-hour target turnaround.

Add about A$24 for 10 GB of logs and around A$1 a month of storage for the first year of retained PDFs, and the Azure bill sits somewhere between roughly A$100 and A$310 a month depending on the path. The cheaper path costs more in engineering and ongoing testing, so the right answer depends on how many invoice layouts you see and who maintains the prompts. For the architecture behind the second path, see our multi-agent document processing architecture guide.

Speech workloads follow the same logic. Transcribing 500 hours of recorded meetings on batch costs A$128.20; the same audio through real-time transcription costs A$712.30. If nobody needs the transcript while the meeting is happening, batch is the default.


Microsoft 365 Copilot and Copilot Studio

Copilot licences come from a different budget line (per user, not per token) but they end up in the same conversation.

  • Microsoft 365 Copilot is listed on Microsoft's Australian site at A$44.90 per user per month paid yearly, or A$47.15 per user per month paid monthly on an annual commitment, excluding GST. It needs a qualifying Microsoft 365 licence underneath. A Copilot Business plan is also listed for organisations licensing up to 300 users; check Microsoft's site for the current price because it has carried promotional pricing.
  • Copilot Studio bills in Copilot Credits. Microsoft's Australian pricing page lists a capacity pack of 25,000 credits at A$299.30 per pack per month, and the Azure pay-as-you-go meter is A$0.0142 per credit (US$0.01). How many credits an agent consumes depends on what it does (a scripted answer, a generative answer, an action, grounding on tenant data), and the rates per feature are set out in Microsoft's Copilot Studio licensing guide.

At list price, 100 Copilot users cost A$4,490 a month, or A$53,880 a year. That is why the order of operations matters: run the Copilot readiness and permissions audit first, license a pilot group, measure, then expand. Our Copilot readiness service covers that sequence. Choosing between Copilot Studio and a custom agent on Azure OpenAI comes down to data, licensing shape and who will maintain it, which is a scoping question more than a pricing one.


A four-week way to get a number you can defend

A spreadsheet estimate rests on guessed prompt sizes and guessed usage. A short pilot replaces both guesses with measurements. This is the sequence we suggest:

From guess to defensible Azure AI budget

1
Week 1
Set the boundaries
Agree the residency rule, the model shortlist and the deployment type. This decides the price list you are using.
2
Week 2
Build a thin pilot
Global or regional Standard, Basic search, logging on. Tag every resource with a cost centre.
3
Week 3
Measure real usage
Pull tokens per question, cache hit rate and search queries from Azure Monitor and Cost Management.
4
Week 4
Model the production bill
Scale measured unit costs to forecast volume, add SLA replicas, non-production and logs, set budget alerts.

A few habits keep the bill honest after go-live: set Azure budget alerts at 50%, 80% and 100% of the monthly forecast, review cost by tag once a month, record which model version each deployment runs so an upgrade does not change the price without anyone noticing, and recheck the price list each quarter. Where nobody internal has time for that, it is the kind of ongoing work our managed AI services cover.


Getting a budget you can take to finance

Azure AI pricing in Australia is transparent once you know which meters you are on. The hard parts are the decisions that come before the arithmetic: which residency boundary you actually need, whether the search tier matches your security design, and who owns the system once it is live.

If you want a second opinion on an estimate, or help turning a pilot into a budget your finance team will accept, book a scoping conversation or contact our team. Our AI strategy work starts with exactly this kind of cost and residency model.


Common Questions

How much does Azure OpenAI cost in Australia?

As checked in October 2026, GPT-5.4 mini on Global Standard in Australia East lists at A$1.0685 per million input tokens and A$6.4107 per million output tokens, excluding GST. Older GPT-4.1 mini on regional Standard is cheaper per token. Confirm current rates in the Azure pricing calculator.

Does Azure OpenAI keep my data in Australia?

Microsoft states that data stored at rest stays in your chosen Azure geography for every deployment type. Prompt processing differs: Global may use any region, Data Zone stays within Asia Pacific, and regional Standard or Regional Provisioned processes within the geography you choose.

When is provisioned throughput cheaper than pay-as-you-go?

Only at high, sustained utilisation. A 15-PTU GPT-5.4 mini reservation costs about A$5,556 a month and roughly matches the pay-as-you-go price only when it runs at full capacity every hour of the month. Business-hours internal tools are usually cheaper on pay-as-you-go.

What is usually the biggest cost in an Azure AI knowledge assistant?

In our hypothetical estimate it was Azure AI Search, at about 65% of the monthly bill, because search bills by the hour and an uptime SLA needs two replicas. Model tokens were about 22%. Check search tier and replica count before optimising prompts.

How much does Microsoft 365 Copilot cost per user in Australia?

Microsoft's Australian site lists Microsoft 365 Copilot at A$44.90 per user per month paid yearly, or A$47.15 paid monthly on an annual commitment, excluding GST. It requires a qualifying Microsoft 365 licence, so run a permissions audit and a pilot before buying at scale.

Is Azure AI Document Intelligence cheaper than using a language model for extraction?

It depends on the workload. Prebuilt models cost A$14.246 per 1,000 pages with little build effort. Read OCR plus a mini model on Batch can cost far less per page, but you build, test and maintain the extraction prompts yourself.


Related Reading:

Sources: Unit prices from the Microsoft Azure Retail Prices API (prices.azure.com, Australia East, AUD), as checked 8 October 2026. Deployment types and data residency from Microsoft Learn, "Understanding deployment types in Microsoft Foundry Models" (August 2026). Model availability from Microsoft Learn, "Region availability for Foundry Models sold by Azure" (September 2026). PTU minimums and throughput from Microsoft Learn, "Determine PTU sizing for a workload" (September 2026) and "Provisioned throughput for Foundry Models" (July 2026). Prompt caching rules from Microsoft Learn, "Prompt caching with Azure OpenAI" (August 2026). AI Search limits from Microsoft Learn, "Service limits for tiers and SKUs" (September 2026). Microsoft 365 Copilot and Copilot Studio AUD prices from microsoft.com/en-au Copilot pricing pages, as checked October 2026. All worked examples are hypothetical.