Technology
Danish Kapoor
Danish Kapoor

Nano Banana 2.1 improves character consistency with 14 reference images

Google has updated its image production and chat editing model with Nano Banana 2.1. The model, which is the sequel to Nano Banana 2, focuses on character consistency, command compliance and the creation of text within the image. The company’s developer document, updated on October 6, lists the new version with the code gemini-nano-banana-2.1.

The model supports using up to 14 reference images in a single production. Besides this, Google explains the capacity to preserve the appearance of four characters and the properties of ten objects. This support gives more guidance to developers who want to use the same character in different scenes; It does not guarantee a perfect match for every result.

On the visual output side, there are 1K, 2K and 4K resolution options; the default option remains 1K. In addition, the model aims to reduce the appearance of repetitive parts in wide and panoramic proportions. Google specifically states the correction of 1:4, 4:1, 1:8 and 8:1 ratios in 2K and 4K outputs.

Nano Banana 2.1 also improves text generation and infographic layout in images. On the other hand, the fact that the model creates better text does not alone indicate that the prepared text is informationally correct. The user continues to further check the suitability of the numbers, labels and descriptions in the image to the source.

Nano Banana 2.1 expands production and editing options

Developers can set the thinking level with minimal, medium or high options. Google uses medium by default. In addition, production supported by web and visual search allows the model to benefit from external information and visual context; This function does not create the same capability as standard file search.

Google’s model card completes the release’s technical framework and evaluation information. The developer document lists text, image, video and PDF input, and image and text output. However, audio production, live API, code execution and function calling are not among the supported capabilities.

The input limit is announced as 131,072 and the output limit is 32,768 tokens. These numbers represent the context and response limits of the model; It does not directly show the number of images that can be produced in the same numbers. In addition, Batch API support offers an option to developers who want to execute batch tasks with a separate processing flow.

Google positions the model as a continuation of the Flash-level speed and cost efficiency approach. However, since the usage cost depends on the selected output and API pricing, a single fixed fee does not result. The new version offers a technical update specifically addressing the need to work with reference and maintain visual integrity across multiple rounds of editing.

Danish Kapoor