Wednesday, October 7, 2026

Google DeepMind Unveils EmbeddingGemma 2 for Local Multimodal AI Search Offline Operation

Google DeepMind Unveils EmbeddingGemma 2 for Local Multimodal AI Search Offline Operation 


Google DeepMind Unveils EmbeddingGemma 2 for Local Multimodal AI Search Offline Operation 


Google DeepMind has introduced EmbeddingGemma 2, an open-weight model engineered to perform privacy-focused, cross-modal semantic searches directly on consumer devices. 


Operating fully offline without relying on cloud processing, the 740-million-parameter model is released under an Apache 2.0 license.


 It maps combinations of text, code, images, video, and audio into a single unified embedding space, enabling instant vector search, zero-shot classification, and low-latency local execution.


Designed for resource-constrained hardware, EmbeddingGemma 2 features a highly efficient modular encoder system that allows devices to load only the necessary components per query. 


Built on the Gemma 4 architecture, it requires approximately 191MB of memory for text-only operations and around 567MB for full multimodal performance on flagship smartphones like the Google Pixel 11 Pro. 


The model also incorporates quantization-aware training, an expanded context window of up to 8,000 tokens, and dynamic vector sizing to minimize local storage footprint.


To showcase its real-world capabilities, Google updated its AI Edge Gallery app with features like "Instant Media Search" and "Video Moments Finder," which allow users to query local media using natural language or image samples. 


Additionally, a new application titled Google AI Edge Foresight for Mac leverages EmbeddingGemma 2 to index system transcripts, audio recordings, and offline files. 


By processing all embeddings and search requests directly on the user's device, these applications ensure that sensitive personal data remains completely private.


Google is backing the launch with broad integration options to help developers deploy on-device semantic retrieval across multiple ecosystems. 


The company plans to deliver EmbeddingGemma 2 as an Android service via ML Kit, offering hardware acceleration out of the box.


 Cross-platform deployment tools provided through LiteRT and MediaPipe Tasks will further extend support across iOS, macOS, Windows, Linux, and web applications, ensuring seamless framework compatibility.


Succeeding an earlier text-focused version that surpassed 20 million downloads, EmbeddingGemma 2 marks a major step forward in multimodal processing efficiency. 


Technical benchmarks demonstrate improved performance across tests like the Massive Text Embedding Benchmark Code and Massive Audio Embedding Benchmark. 


Developers can now access model weights and deployment resources through open AI hubs including Hugging Face and Kaggle.



Google DeepMind ने लोकल मल्टीमॉडल AI सर्च ऑफ़लाइन ऑपरेशन के लिए EmbeddingGemma 2 पेश किया

Google DeepMind ने EmbeddingGemma 2 पेश किया है, जो एक ओपन-वेट मॉडल है जिसे सीधे कंज्यूमर डिवाइस पर प्राइवेसी-फोकस्ड, क्रॉस-मॉडल सिमेंटिक सर्च करने के लिए बनाया गया है।

क्लाउड प्रोसेसिंग पर निर्भर हुए बिना पूरी तरह ऑफ़लाइन काम करने वाला, 740-मिलियन-पैरामीटर मॉडल Apache 2.0 लाइसेंस के तहत रिलीज़ किया गया है।

यह टेक्स्ट, कोड, इमेज, वीडियो और ऑडियो के कॉम्बिनेशन को एक सिंगल यूनिफाइड एम्बेडिंग स्पेस में मैप करता है, जिससे इंस्टेंट वेक्टर सर्च, ज़ीरो-शॉट क्लासिफिकेशन और लो-लेटेंसी लोकल एग्जीक्यूशन मुमकिन होता है।

रिसोर्स-कंस्ट्रेंड हार्डवेयर के लिए डिज़ाइन किया गया, EmbeddingGemma 2 में एक बहुत एफिशिएंट मॉड्यूलर एन्कोडर सिस्टम है जो डिवाइस को हर क्वेरी के लिए सिर्फ़ ज़रूरी कंपोनेंट लोड करने देता है।

Gemma 4 आर्किटेक्चर पर बना, इसे सिर्फ़ टेक्स्ट ऑपरेशन के लिए लगभग 191MB मेमोरी और Google Pixel 11 Pro जैसे फ्लैगशिप स्मार्टफ़ोन पर पूरी मल्टीमॉडल परफॉर्मेंस के लिए लगभग 567MB मेमोरी की ज़रूरत होती है। 

इस मॉडल में क्वांटाइज़ेशन-अवेयर ट्रेनिंग, 8,000 टोकन तक की एक एक्सपैंडेड कॉन्टेक्स्ट विंडो और लोकल स्टोरेज फुटप्रिंट को कम करने के लिए डायनामिक वेक्टर साइज़िंग भी शामिल है।

अपनी रियल-वर्ल्ड कैपेबिलिटीज़ को दिखाने के लिए, Google ने अपने AI Edge Gallery ऐप को "Instant Media Search" और "Video Moments Finder" जैसे फीचर्स के साथ अपडेट किया है, जो यूज़र्स को नेचुरल लैंग्वेज या इमेज सैंपल्स का इस्तेमाल करके लोकल मीडिया को क्वेरी करने की सुविधा देते हैं।

इसके अलावा, Mac के लिए Google AI Edge Foresight नाम का एक नया एप्लिकेशन सिस्टम ट्रांसक्रिप्ट, ऑडियो रिकॉर्डिंग और ऑफलाइन फाइलों को इंडेक्स करने के लिए EmbeddingGemma 2 का इस्तेमाल करता है।

यूज़र के डिवाइस पर सीधे सभी एम्बेडिंग और सर्च रिक्वेस्ट को प्रोसेस करके, ये एप्लिकेशन यह पक्का करते हैं कि सेंसिटिव पर्सनल डेटा पूरी तरह से प्राइवेट रहे।

Google डेवलपर्स को कई इकोसिस्टम में ऑन-डिवाइस सिमेंटिक रिट्रीवल डिप्लॉय करने में मदद करने के लिए बड़े इंटीग्रेशन ऑप्शन के साथ लॉन्च को सपोर्ट कर रहा है।

कंपनी ML Kit के ज़रिए एक Android सर्विस के तौर पर EmbeddingGemma 2 देने का प्लान बना रही है, जो आउट ऑफ द बॉक्स हार्डवेयर एक्सेलरेशन ऑफर करता है।

 LiteRT और MediaPipe Tasks के ज़रिए दिए गए क्रॉस-प्लेटफ़ॉर्म डिप्लॉयमेंट टूल iOS, macOS, Windows, Linux और वेब एप्लिकेशन में सपोर्ट को और बढ़ाएंगे, जिससे फ्रेमवर्क कम्पैटिबिलिटी आसान हो जाएगी।

पहले के टेक्स्ट-फ़ोकस्ड वर्शन, जिसके 20 मिलियन से ज़्यादा डाउनलोड हुए थे, के बाद EmbeddingGemma 2 मल्टीमॉडल प्रोसेसिंग एफ़िशिएंसी में एक बड़ा कदम है।

टेक्निकल बेंचमार्क मैसिव टेक्स्ट एम्बेडिंग बेंचमार्क कोड और मैसिव ऑडियो एम्बेडिंग बेंचमार्क जैसे टेस्ट में बेहतर परफ़ॉर्मेंस दिखाते हैं।

डेवलपर्स अब हगिंग फ़ेस और कैगल जैसे ओपन AI हब के ज़रिए मॉडल वेट और डिप्लॉयमेंट रिसोर्स एक्सेस कर सकते हैं।

ఆఫ్‌లైన్‌లో స్థానిక మల్టీమోడల్ AI శోధన కోసం గూగుల్ డీప్‌మైండ్ EmbeddingGemma 2ను ఆవిష్కరించింది

గూగుల్ డీప్‌మైండ్, వినియోగదారుల పరికరాల్లోనే నేరుగా గోప్యత-కేంద్రీకృత, క్రాస్-మోడల్ సెమాంటిక్ శోధనలను నిర్వహించడానికి రూపొందించిన ఓపెన్-వెయిట్ మోడల్ అయిన EmbeddingGemma 2ను పరిచయం చేసింది.

క్లౌడ్ ప్రాసెసింగ్‌పై ఆధారపడకుండా పూర్తిగా ఆఫ్‌లైన్‌లో పనిచేసే ఈ 740-మిలియన్-పారామీటర్ల మోడల్, అపాచీ 2.0 లైసెన్స్ క్రింద విడుదల చేయబడింది.

ఇది టెక్స్ట్, కోడ్, చిత్రాలు, వీడియో మరియు ఆడియోల కలయికలను ఒకే ఏకీకృత ఎంబెడింగ్ స్పేస్‌లోకి మ్యాప్ చేస్తుంది. తద్వారా తక్షణ వెక్టర్ శోధన, జీరో-షాట్ వర్గీకరణ మరియు తక్కువ-లేటెన్సీతో కూడిన స్థానిక అమలును ఇది సాధ్యం చేస్తుంది.

పరిమిత వనరులు గల హార్డ్‌వేర్ కోసం రూపొందించబడిన EmbeddingGemma 2, అత్యంత సమర్థవంతమైన మాడ్యులర్ ఎన్‌కోడర్ సిస్టమ్‌ను కలిగి ఉంది. ఇది ప్రతి క్వెరీకి అవసరమైన భాగాలను మాత్రమే పరికరాలు లోడ్ చేసుకోవడానికి అనుమతిస్తుంది.

 జెమ్మా 4 ఆర్కిటెక్చర్‌పై నిర్మించబడిన దీనికి, గూగుల్ పిక్సెల్ 11 ప్రో వంటి ఫ్లాగ్‌షిప్ స్మార్ట్‌ఫోన్‌లలో టెక్స్ట్-మాత్రమే కార్యకలాపాల కోసం సుమారు 191MB మెమరీ మరియు పూర్తి మల్టీమోడల్ పనితీరు కోసం దాదాపు 567MB మెమరీ అవసరం.

ఈ మోడల్‌లో క్వాంటైజేషన్-అవేర్ ట్రైనింగ్, 8,000 టోకెన్‌ల వరకు విస్తరించిన కాంటెక్స్ట్ విండో, మరియు లోకల్ స్టోరేజ్ ఫుట్‌ప్రింట్‌ను తగ్గించడానికి డైనమిక్ వెక్టర్ సైజింగ్ వంటివి కూడా పొందుపరచబడ్డాయి.

దాని వాస్తవ-ప్రపంచ సామర్థ్యాలను ప్రదర్శించడానికి, గూగుల్ తన AI ఎడ్జ్ గ్యాలరీ యాప్‌ను "ఇన్‌స్టంట్ మీడియా సెర్చ్" మరియు "వీడియో మూమెంట్స్ ఫైండర్" వంటి ఫీచర్లతో అప్‌డేట్ చేసింది. ఇవి వినియోగదారులను సహజ భాష లేదా చిత్ర నమూనాలను ఉపయోగించి స్థానిక మీడియాను శోధించడానికి అనుమతిస్తాయి.

అదనంగా, Mac కోసం గూగుల్ AI ఎడ్జ్ ఫోర్‌సైట్ అనే కొత్త అప్లికేషన్, సిస్టమ్ ట్రాన్స్‌క్రిప్ట్‌లు, ఆడియో రికార్డింగ్‌లు మరియు ఆఫ్‌లైన్ ఫైల్‌లను ఇండెక్స్ చేయడానికి ఎంబెడింగ్‌జెమ్మా 2ను ఉపయోగించుకుంటుంది.

వినియోగదారు పరికరంలోనే అన్ని ఎంబెడింగ్‌లు మరియు శోధన అభ్యర్థనలను నేరుగా ప్రాసెస్ చేయడం ద్వారా, ఈ అప్లికేషన్‌లు సున్నితమైన వ్యక్తిగత డేటా పూర్తిగా గోప్యంగా ఉండేలా చూస్తాయి.

 బహుళ ఎకోసిస్టమ్‌లలో డెవలపర్‌లు ఆన్-డివైస్ సెమాంటిక్ రిట్రీవల్‌ను అమలు చేయడానికి సహాయపడే విస్తృతమైన ఇంటిగ్రేషన్ ఎంపికలతో గూగుల్ ఈ లాంచ్‌కు మద్దతు ఇస్తోంది.

కంపెనీ, ML కిట్ ద్వారా ఎంబెడింగ్‌జెమ్మా 2ను ఒక ఆండ్రాయిడ్ సర్వీస్‌గా అందించాలని యోచిస్తోంది, ఇది అవుట్-ఆఫ్-ది-బాక్స్ హార్డ్‌వేర్ యాక్సిలరేషన్‌ను అందిస్తుంది.

లైట్‌ఆర్‌టి మరియు మీడియాపైప్ టాస్క్‌ల ద్వారా అందించబడిన క్రాస్-ప్లాట్‌ఫామ్ డిప్లాయ్‌మెంట్ టూల్స్, ఐఓఎస్, మాక్‌ఓఎస్, విండోస్, లైనక్స్ మరియు వెబ్ అప్లికేషన్‌లలో మద్దతును మరింత విస్తరిస్తాయి, తద్వారా ఫ్రేమ్‌వర్క్ అనుకూలత సజావుగా ఉండేలా చూస్తాయి.

20 మిలియన్ల డౌన్‌లోడ్‌లను అధిగమించిన మునుపటి టెక్స్ట్-కేంద్రీకృత వెర్షన్ తర్వాత వచ్చిన ఎంబెడింగ్‌జెమ్మా 2, మల్టీమోడల్ ప్రాసెసింగ్ సామర్థ్యంలో ఒక పెద్ద ముందడుగును సూచిస్తుంది.

మాసివ్ టెక్స్ట్ ఎంబెడింగ్ బెంచ్‌మార్క్ కోడ్ మరియు మాసివ్ ఆడియో ఎంబెడింగ్ బెంచ్‌మార్క్ వంటి పరీక్షలలో సాంకేతిక బెంచ్‌మార్క్‌లు మెరుగైన పనితీరును ప్రదర్శించాయి.

డెవలపర్‌లు ఇప్పుడు హగ్గింగ్ ఫేస్ మరియు కాగిల్ వంటి ఓపెన్ ఏఐ హబ్‌ల ద్వారా మోడల్ వెయిట్స్ మరియు డిప్లాయ్‌మెంట్ వనరులను యాక్సెస్ చేయవచ్చు.

No comments:

Post a Comment

Please Dont Leave Me