Google Gemini

Google Gemini is a family of advanced multimodal artificial intelligence (AI) models developed by Google. Introduced as the successor to Google Bard, Gemini is designed to be an AI assistant capable of understanding and processing various types of input, such as text, images, audio, video, and programming code. Since its official launch in February 2024, Gemini has evolved into one of the most comprehensive AI ecosystems, deeply integrated with Google services and available in various subscription tiers to meet the needs of users ranging from students to enterprises.

History and Development

The roots of Google Gemini lie in Google Bard, a chatbot based on the LaMDA language model that was launched in 2023. A major transformation took place in early 2024 when Bard was upgraded with a more advanced next-generation large language model (LLM). On February 8, 2024, Google officially announced the name change from Bard to Gemini, marking a strategic shift to unify all of its AI capabilities under a single brand umbrella. This change was also accompanied by the launch of the Gemini app for smartphones, which gradually replaced Google Assistant as the primary digital assistant on certain Android devices.

Throughout 2025 and 2026, Google continued to develop Gemini models at a high pace. The launch of models such as Gemini 3.5 Flash and Gemini 3.1 Flash-Lite demonstrated Google's commitment to providing models with a balance between high intelligence and low latency, especially for programming and agentic tasks. Recent innovations include the introduction of Gemini Omni Flash, a multimodal model capable of generating 720p-quality video from text descriptions or editing video through simple conversation. This model was released in public preview in late June 2026, marking a major step for Google in the field of generative content creation.

Core Capabilities and Features

Multimodality and Content Processing

Gemini's main feature is its capability as a multimodal model. Users can interact with Gemini through text, voice, and images. The Gemini app allows users to take photos and ask questions about objects around them. This capability is extended with the video-to-image feature on the latest model, in which Gemini can generate high-quality thumbnails or posters from uploaded video files. Gemini is also able to create Audio Overviews, turning documents and reports into engaging podcast-style discussions, making complex content easier to consume.

Integration and Extensions

Gemini is designed to work seamlessly with the Google ecosystem. Through the extensions feature, Gemini can connect with apps such as Gmail, Google Maps, YouTube, and Google Drive to help complete tasks more quickly. This capability allows Gemini, for example, to summarize email contents, search for travel routes, or find information from videos without having to open the apps separately. On Pixel 9 devices and later, Gemini is the default assistant, accessible by long-pressing the power button or using the voice command "Hey Google".

Personalization with Gems

One of the main personalization features is Gems. Gems are specialized versions of Gemini that can be customized to become "experts" on a particular topic without requiring coding skills. Users can create Gems for specific purposes such as writing editors, event planners, or sentiment analysis, by uploading instructions and reference files. Google also provides ready-made Gems, such as Learning coach for study guidance and Career guide for career preparation, which are highly useful in the education sector.

Gemini Live and Conversation

The Gemini Live feature delivers a more natural conversational experience. Users can speak with Gemini in real time, use it to brainstorm ideas, or simulate important conversations. Gemini Live is designed for smooth interaction, allowing users to share their screen or camera while speaking in order to receive contextual responses.

Models and Versions

Google releases Gemini in various model variants to meet different needs and budgets. Each model has optimized specifications, ranging from speed and cost to intelligence and multimodal capability.

Gemini 3.5 Flash

This model is a mainstay for tasks that require substantial intelligence at high speed. Released generally in May 2026, Gemini 3.5 Flash is known to be highly effective for programming tasks and long-term "agentic" tasks, offering performance comparable to large "flagship" models at a lower cost.

Gemini 3.1 Flash-Lite

This variant is optimized for speed, scale, and cost efficiency. The model is well suited to simple tasks and image generation with ultra-low latency.

Gemini Omni Flash

The newest model, still in public preview. Its main capability is conversational video creation and editing. Equipped with an invisible SynthID digital watermark for content verification, this model represents a new generation of generative AI capable of creating content from various types of input.

Image Models (Nano Banana)

Gemini also has specialized visual models such as Gemini 3.1 Flash Image ("Nano Banana 2") and Gemini 3 Pro Image ("Nano Banana Pro"). These models have been generally available and excel at image creation and editing, including a feature for generating images from video.

Availability and Subscriptions

Gemini can be accessed through various platforms, including mobile apps for Android (Android 10+) and iOS, as well as through the website at gemini.google.com. Google offers several subscription plans.

The Google AI Pro plan unlocks access to the most advanced models such as 2.5 Pro, the Deep Research feature for detailed reports, and eight-second video generation with Veo 3, as well as a 1-million-token context window capable of processing up to 1,500 pages of text.

The Google AI Ultra plan is the highest tier, offering exclusive access to the most powerful models, the latest features such as Agent Mode, and the opportunity to try Google's AI innovations earlier. This plan is available in certain countries and is not for Workspace customers.

In Southeast Asia, Gemini has recorded significant user growth, more than doubling in a year. Vietnam has become the regional leader in using Gemini for academic purposes and local-language use. The country has even become the market with the highest percentage of native-language (Vietnamese) use in the region, reaching 89%.

Impact and Education

Gemini, along with other generative AI, has begun to change the way people access information. This change has sparked discussion about the future of traditional knowledge sources such as Wikipedia. Wikimedia Indonesia, for example, held the #WikipediaxAI program in 2025 to discuss how AI uses Wikipedia content and how the community can adapt. Concerns have arisen that the shift from reading articles to receiving instant summaries from AI could reduce the number of visitors to source sites and the number of new volunteer contributors.

On the other hand, Gemini offers great potential in education. In Vietnam, more than 160,000 students use the Gemini Canvas feature every month for exam preparation, while educators generate tens of thousands of teaching assistance requests every day. This use shows AI's role as a powerful supporting tool, not a replacement for human critical thinking. A comparative study even placed Gemini as a tool comparable to ChatGPT for various translation tasks, demonstrating its capability in language processing.

References

1. Use Gemini, your AI assistant, on your Pixel phone. Google Help.

2. Education/News/December 2025/WikipediaxAI: Wikipedia, AI, and the future of knowledge. Meta-Wiki, Wikimedia.

3. Catatan rilis, Gemini API. Google AI for Developers.

4. Personalize Gemini with Gems: custom AI experts for any topic. Google Services.

5. Google Gemini, Apps on Google Play.

6. Cara Menggunakan Gemini AI Google Bahasa Indonesia, Mudah dan Praktis. Kompas.com.

7. Release notes, Gemini API. Google AI for Developers.

8. Gemini, "Asisten" Canggih untuk Sehari-hari. LinkedIn, 360 Solusi Teknologi.

9. Việt Nam dẫn đầu Đông Nam Á về dùng AI của Google hỗ trợ học thuật. VietnamPlus.

10. Digilib UIN SGD (Bab 1).

11. 100 things we announced at I/O 2026. Google Blog.

12. Try Gemini, your personal AI assistant. Android.com.

[Read more →]

Nano Banana

Nano Banana is the code name for an advanced artificial intelligence (AI)–based image generation and editing model developed by Google. Introduced in late August 2025 as part of the Gemini model family, specifically Gemini 2.5 Flash Image, the technology quickly became a viral phenomenon on social media because of its ability to generate and edit images with high quality, precision, and exceptional character consistency using only natural language prompts. The name "Nano Banana" itself comes from an efficient Gemini model variant ("Nano") and an unusual internal nickname during development ("Banana").

History and Development

Speculation about the existence of "Nano Banana" began circulating among internet users after Google engineers and eventually CEO Sundar Pichai posted banana emojis on social media. The mystery was answered on August 26, 2025, when Google officially announced the launch of its latest image model, Gemini 2.5 Flash Image, whose code name is Nano Banana. This launch marked a significant improvement over the previous generation image model, Gemini 2.0 Flash, with a focus on higher image quality and stronger creative control for users.

Since its launch, Nano Banana has experienced a spectacular surge in popularity. In its first month alone, more than 500 million images were generated by users around the world using its visual editing feature. Thanks to this model, the Gemini app succeeded in attracting 23 million new users and rose to the top of the App Store and Google Play in various countries, including Indonesia.

Core Capabilities and Features

Image Generation and Editing with High Consistency

Nano Banana's main advantage lies in its ability to produce unmatched image consistency. Unlike many other AI generators that often struggle to maintain visual continuity, Nano Banana excels at preserving intricate details such as facial expressions, hand movements, lighting, textures, and spatial relationships across generated elements. This capability is especially important for projects that require a uniform appearance of characters, style, or environmental details.

Users can upload an existing image and give natural language text instructions to edit it, such as changing the background, replacing clothing, or adding objects, while the model ensures that the main subject remains recognizable and consistent. For example, someone can upload a photo of themselves and see how they would look with a different haircut or particular clothing, all while still looking like themselves.

Natural Language Understanding and Conversational Editing

Nano Banana is designed to understand complex instructions in everyday language. Users do not need technical expertise such as using Photoshop; they simply describe what they want, for example "replace the cloudy sky with a sunset and add a flying drone," and the model will execute it. The conversational editing feature allows users to carry out a series of edits in sequence while keeping the image style and lighting consistent.

Multi-Image Merging

Another innovative capability is multi-image merging, which allows users to combine several images into one cohesive visual. Google gave the example that someone can upload a photo of themselves and a photo of their dog to create "the perfect portrait of the two of you on a basketball court." This fusion technology creates a perfect composition with a simple prompt.

Models in the Nano Banana Family

The platform offers several model variants tailored to different needs, all built on Gemini's multimodal capabilities and responsive even to the most detailed prompts.

Nano Banana Pro is the flagship model designed for image generation and editing with professional studio-level precision and control. This model is built on Gemini 3 and allows users to imagine and create almost anything, from detailed images to realistic or imaginative ones. Nano Banana 2 offers professional-level image generation and editing at lightning speed, based on the Gemini 3.1 Flash Image model. Meanwhile, Nano Banana 2 Lite is the fastest and most efficient model, providing the highest speed and lowest cost for generation and editing.

Impact and Viral Phenomenon

3D Action Figurine Trend

Nano Banana became highly viral thanks to the trend of turning ordinary photos into highly realistic 3D action figurines. Users used prompts to turn self-portraits, celebrities, or cartoon characters into collectible statues that look like toys in packaging, often with bases and custom packaging that make them look like products on a store shelf. This trend flooded social media platforms such as X, Reddit, and Instagram.

Democratization of Aesthetics and Democratization of Creativity

The Nano Banana phenomenon marks a kind of democratization of aesthetics, in which anyone can look like a celebrity or create professional-class visuals without needing design skills or expensive equipment. In Indonesia, this culture was quickly adopted, where people from various backgrounds can turn simple photos into stunning works of art.

Cross-Platform Integration

Google has aggressively integrated Nano Banana into its application ecosystem. After its success in the Gemini App and Google AI Studio, the model is in the process of being integrated into Google Messages, allowing users to edit and generate images directly from text conversations. It has also been demonstrated in Google Lens and is being developed for Google Photos.

Access and Availability

Nano Banana can be accessed globally through the Gemini app or Google AI Studio. The model is available to free and paid users, although free users may have daily production limits and slower processing speeds, while paid accounts remove most of those limitations.

All images created or edited with Nano Banana include a SynthID digital watermark that is both visible and invisible to indicate that the image was generated by AI.

Impact and Criticism

Although it offers unlimited ease and creativity, the emergence of Nano Banana has also raised questions about authenticity and psychological impact. The ability to drastically change appearance so easily can lead to cultural dissonance, in which people trust their digital selves more than their real selves. There are concerns that a generation of users may become more occupied with arranging AI prompts than living real life.

The question of "who we are without filters, without edits, without a nano banana overlaying our faces" has become an important reflection amid the onslaught of technology that increasingly blurs the line between reality and virtuality.

References

Google DeepMind. "Nano Banana". Google DeepMind. Diakses pada 15 Juli 2026.

Media.io. "Nano Banana AI: Apakah Ini Akhir dari Pengeditan Gambar Tradisional? Inilah Cara Menguasainya". Media.io. 6 September 2025.

Hindustan Times. "'Nano Banana' is taking over internet: All about viral AI action figure craze". Hindustan Times. 11 September 2025.

Trusted Reviews. "What is Nano Banana? Google’s mysterious technology is finally unveiled". Trusted Reviews. 28 Agustus 2025.

Duta.co. "Pribadi Nano Banana". Duta.co Berita Harian Terkini. 18 September 2025.

Radar Banyuwangi. "Viral! Tren Sulap Foto Lama Jadi Kekinian Pakai Gemini AI, Gampang Banget!". Radar Banyuwangi. 25 September 2025.

Alibaba.com Reads. "Bersedia untuk Kejutan 'Pisang Nano' dalam Mesej Google". Alibaba.com Reads. 19 Oktober 2025.

TEGAROOM ONLINE. "Mengubah Imajinasi Menjadi Visual: Google Perkenalkan Nano Banana 2025". 30 November 2025.

[Read more →]