Per-image billing at three resolutions
$0.134 for 1K or 2K and $0.24 for 4K, against $0.067 and $0.12 on batch. Input runs $2.00 per million tokens, about $0.0011 per input image, and text and thinking output $12.00 per million.
Google's top image model prices per image rather than per token: $0.134 for 1K or 2K, $0.24 at 4K, and half that on batch. For anyone building a fixed-price product or writing a proposal, that legibility is worth more than a marginal quality difference. It also renders legible stylised text, which is the specific failure that used to keep generated imagery out of commercial work.
$0.134 for a 1K or 2K image, $0.24 at 4K, and roughly half those on batch. You can build a cost model, write a proposal, and set a plan limit before you have generated anything. OpenAI's GPT-Image-2 bills at $30 per million output tokens, which means measuring your own prompts before you can quote at all.
Legible stylised text for infographics, menus, diagrams, and marketing assets. Mangled lettering was the single most common reason generated imagery got rejected in commercial review, and it turned every asset into a manual fix. This is a bigger practical change than another increment of photorealism.
Nano Banana 2 costs half as much at 1K and supports up to 10 object references, 4 character references for consistency, and 3 style references. Pro takes 6 object references and supports neither character consistency nor style references. For brand work, test the cheaper model first, which is not the sentence anyone expects to write.
SynthID goes on everything the model produces. For a compliance position that requires provenance marking on synthetic media, that is a reason to choose it. For a client who does not want their marketing assets identifiable as generated, it is a reason to have the conversation early rather than late.
$0.067 per 1K or 2K image and $0.12 at 4K on the batch path. Catalogue imagery, variants, and backfills do not need a synchronous response, and moving them is the cheapest optimisation available here.
Two comparisons decide the build: against the cheaper Google model that has more features, and against OpenAI's token-priced alternative.
| Dimension | vs Nano Banana 2 | vs GPT-Image-2 |
|---|---|---|
Cost per image | Roughly double at every shared resolution: $0.134 against $0.067 at 1K, $0.24 against $0.151 at 4K. Nano Banana 2 also offers a 512px tier at $0.045 that Pro does not, which matters for thumbnails and previews. | A knowable number against a measured average. GPT-Image-2 bills output at $30 per million tokens with cost rising with resolution and detail, so a fixed-price engagement has to run a measurement exercise before it can commit to anything. |
Reference images | Pro takes 6 object references and does not support character consistency or style references. Nano Banana 2 takes 10 object, 4 character, and 3 style references. On this axis the cheaper model is simply more capable. | Both Google models expose an explicit, documented reference surface with counts you can design around. That specificity is what makes a repeatable brand pipeline possible rather than a prompt-tuning exercise. |
Output control | Nano Banana 2 documents ten aspect ratios from 21:9 to 9:16 and four resolution tiers including 512px. If you are generating for several placements from one brief, that surface saves a cropping stage. | Google publishes resolutions and aspect ratios as a spec. On the OpenAI side these are handled through the request rather than as a documented tier list, which makes capacity planning less direct. |
Provenance | SynthID on both Google models, so this is not a differentiator between them. It is a differentiator when your compliance position requires a marker and you are choosing across vendors. | Confirm what OpenAI applies for your use rather than assuming parity. Provenance policy is moving quickly under the EU AI Act, and this is exactly the sort of thing that changes between releases without a headline. |
Image generation sits behind the same abstraction as every other model call in a Kensink build, so a catalogue pipeline can run on the cheapest model that passes review while a hero asset routes to the top tier, and neither decision is welded into the application.
$0.134 for 1K or 2K and $0.24 for 4K, against $0.067 and $0.12 on batch. Input runs $2.00 per million tokens, about $0.0011 per input image, and text and thinking output $12.00 per million.
Infographics, menus, diagrams, and marketing assets with text that survives review. This removes the manual fix stage that made generated imagery a false economy in a lot of commercial pipelines.
Include specific objects in the output rather than describing them and hoping. Note the boundary: object references yes, character consistency and style references no, and both of those live on Nano Banana 2 instead.
Inpainting that operates on what a region contains rather than on a rectangle you drew. Changing an object in an approved image is the single most common production edit, and this is the right shape for it.
Provenance marking applied automatically, with no changes to your request. Whether that is a feature or a constraint depends entirely on your client's position, which is a conversation to have during scoping.
$0.24 per image standard and $0.12 on batch. Print, large-format, and hero placements no longer need an upscaling stage bolted on afterwards.
The limits, the endpoints, and the lines that decide whether this model fits your transport and your budget. Kept here so nobody has to reconstruct it from three vendor pages.
| Model ID | gemini-3-pro-image |
|---|---|
| Resolutions | 1K, 2K, 4K |
| Reference images | Up to 6 objects at high fidelity |
| Character consistency | Not supported. Nano Banana 2 supports up to 4 character references |
| Style references | Not supported. Nano Banana 2 supports up to 3 |
| Editing | Inpainting with semantic masking |
| Text rendering | Legible stylised text for infographics, menus, diagrams, and marketing assets |
| Provenance | SynthID watermark on every generated image |
| Input pricing | $2.00 / MTok text and image, about $0.0011 per input image |
| Text and thinking output | $12.00 / MTok |
| Image output | $120.00 / MTok, or $0.134 per 1K/2K image and $0.24 per 4K image |
| Batch | $0.067 per 1K/2K image, $0.12 per 4K image, $1.00 / MTok text input, $6.00 / MTok text output |
| Cheaper sibling | Nano Banana 2 (gemini-3.1-flash-image): $0.045 at 512px, $0.067 at 1K, $0.101 at 2K, $0.151 at 4K |
Per image and per million tokens, set against the cheaper Google tier and the OpenAI alternative.
| 1K or 2K image | $0.134 | $0.067 on batch. |
|---|---|---|
| 4K image | $0.24 | $0.12 on batch. |
| Input | $2.00 / MTok | Text and image. About $0.0011 per input image. $1.00 text on batch. |
| Text and thinking output | $12.00 / MTok | $6.00 on batch. |
| Nano Banana 2 | $0.067 / image | 1K. $0.045 at 512px, $0.101 at 2K, $0.151 at 4K. |
| GPT-Image-2 | $30 / MTok out | Token-priced. Image input $8 / MTok, cached $2. |
The reason we usually land on Google for volume image work is not quality, it is that the finance conversation is short. A per-image rate lets you set a plan allowance, price a fixed-scope engagement, and alert on unit cost from the first week. Token-priced generation needs a measurement phase before any of that is possible, and on a fixed-price build that phase is real money. Where OpenAI is already in the stack and the volume is modest, that calculus flips.
Every image the model generates carries the watermark. If your compliance position requires provenance marking on synthetic media, that is a reason to pick this model. If a client expects marketing assets that are not identifiable as generated, it is a constraint they need to know about during scoping rather than after the first delivery.
Ownership of output, the provenance of training data, and exposure from a generated likeness or a recognisable style are unsettled and vary by jurisdiction. We surface them at the start of an engagement, because a brand or legal review is a very expensive place to discover them.
An image model that shifts its defaults will quietly stop matching the assets you already shipped. Keep a reference set of approved output, compare against it when a version moves, and treat aesthetic drift as a regression rather than as a taste question, because no test in your suite will flag it.
Most of our image work is inside a fixed-scope build, and a per-image rate means the cost model is a spreadsheet rather than a research task. That has decided more vendor choices for us than output quality has, which is not what anyone expects going in.
Nano Banana 2 is half the price and has the wider reference surface, including the character consistency Pro lacks. On brand work it has repeatedly been the better answer. We only move up to Pro when a review actually rejects the cheaper output.
Generated imagery with mangled text meant a manual fix on every asset, which erased the saving. Legible text moves whole categories, infographics and menus and merchandising, from novelty into something you can put in a pipeline.
Half price for asynchronous generation, and almost all asset work is asynchronous once examined honestly. We treat a synchronous generation call as the exception that needs a reason.
SynthID is not optional, and clients have strong and opposite reactions to it. Finding out which reaction you are dealing with after the first delivery is a bad week that a single scoping question prevents.
Every model we integrate runs through the same operating system. Three pillars, sixteen layers, one Compound Growth Loop. The methodology that keeps AI work from rotting after the first ship.
Read the K-FrameworkDirect API integration with the model. No LangChain, no orchestration vendor, no agent framework built on quicksand. Typed contracts, the same way we wire up Postgres.
An eval suite built from your real tasks gates every prompt and model change. Quality is measured before it ships, not vibed in a demo.
Governance, audit, and oversight wired in from day one. Who called what, with which prompt version, at what cost. Your auditors get answers, not screenshots.
A model in production without observability is roulette. We instrument every integration so engineering and finance can see the same numbers, and so a regression at 3am surfaces before a customer opens a ticket.
Tokens in, tokens out, dollars spent. Sliced by feature, tenant, and route. Budgets enforced where it matters.
Real distributions, not averages. We know which routes are slow, and why.
The same eval suite that gates a release runs continuously in production. A regression on real traffic surfaces fast.
PII scrubbed at the proxy, shipped to your SIEM. Retention controls match your compliance window.
Dashboards your team owns, not ours. At handoff you get the queries, the alerts, and the runbook. We are not in the path to read your metrics.