Google has released Google AI Edge Foresight, an experimental Mac app that listens to a meeting, expands your rough shorthand into fuller notes, and lets you search past transcripts and files, with all inference running on your machine. Google says it works offline and has no cloud subscription cost. It is the first consumer-facing app built on EmbeddingGemma 2, the open multimodal embedding model Google DeepMind shipped two days ago.
The details come from Google's developer blog post on EmbeddingGemma 2 and Google AI Edge. This post separates what Google has stated from what you still need to verify yourself.
Quick answers
| Question | Answer |
|---|---|
| What is it? | An experimental macOS meeting companion from the Google AI Edge team |
| What does it do? | Enriches your shorthand notes using the live conversation, and searches transcripts, notes, images and documents |
| Which models? | EmbeddingGemma 2 (740M parameters) and Gemma 4 models, per Google |
| Does it need the internet? | Google says no, it works completely offline |
| Does it cost a subscription? | Google says there are no cloud subscription costs |
| Which meeting tools work? | Any, because it reads system audio and the microphone |
| Is it production ready? | Google calls it experimental |
What Foresight actually does
Google's post describes three jobs. The first is enhanced note-taking: you type short fragments during a call and Foresight fills them out with details retrieved in real time from the conversation itself. If you jot "pricing, Q4, Dana" it can attach what was actually said about pricing.
The second is live question answering. Google's demo captions mention "questions detected and answered live by AI," meaning the app notices when someone asks a question in the meeting and drafts an answer from your indexed material.
The third is cross-modal retrieval. Because EmbeddingGemma 2 maps text, audio, images and video frames into a single vector space, Foresight can search images, documents, transcripts and notes with one natural-language query. Google says this runs against your personal knowledge library without leaving the device.
The capture method matters. Foresight integrates directly with system audio and the microphone rather than joining your call as a bot participant. That is why Google says it works "out-of-the-box with any meeting platform." It also means other participants never see a notetaker in the attendee list, which is both convenient and a consent question, covered below.
Why EmbeddingGemma 2 is the interesting part
Most local meeting tools chain several models: speech-to-text, a text embedder for search, maybe a captioner for screenshots. Google's pitch for EmbeddingGemma 2 is that a single compact model replaces that chain. According to the same developer post, the full multimodal model needs about 567MB of active RAM on a Pixel 11 Pro, and the text-only weights about 191MB. On a MacBook M5 Pro GPU, visual embeddings take as little as 37.3 ms per image (roughly 26.9 images per second), measured with a budget of 70 vision tokens per image.
Our EmbeddingGemma 2 explainer covers the model sizes, the Apache 2.0 license and the Matryoshka truncation options in detail. The short version for this story: Foresight is a reference implementation showing what a private, always-available retrieval layer feels like when the embedding model fits comfortably next to a chat model on a laptop.
Google also says EmbeddingGemma 2 can act as a zero-shot decision engine, matching inputs against label descriptions in milliseconds, and the same post shows a MediaPipe Decision task evaluating 500 options per turn in under 100 ms. That is a developer feature rather than something Foresight advertises, but it hints at where the on-device stack is heading.
How it fits Google's local AI push
Foresight did not appear from nowhere. Google has been steadily building out a local stack around Gemma 4:
- Gemma 4 12B brought a multimodal model that fits laptops with 16GB of memory.
- The July Gemma 4 updates improved tool calling and attention speed.
- The Google AI Edge Gallery app, which now gains Instant Media Search and Video Moments Finder demos powered by EmbeddingGemma 2, is the playground for these models on Android and iOS.
Foresight is the Mac counterpart: where the Gallery shows what a model can do, Foresight wraps it into a task people already perform every day. Google says its inference engine, LiteRT, powers ML Kit, MediaPipe Tasks, the Gallery and Foresight, so the same optimizations carry across.
How it compares with other local meeting tools
Foresight lands in a space where open-source projects have been working for months. Meetily is a privacy-first meeting assistant with local transcription. FluidVoice handles on-device dictation on macOS, and Files.md takes the local-first, plain-markdown route to notes.
What differs is the vendor. Foresight is closed-source from Google, though the models underneath (the Gemma family and EmbeddingGemma 2) are open-weight. If you want to inspect or modify the pipeline, the open projects still win. If you want something polished from the team that trains the models, Foresight is the new option. Google has not said whether the app's source will be released, so do not assume it.
What to verify before trusting it with real meetings
Privacy claims are the vendor's. Google says processing is fully local and that sensitive information remains secure on your device. That is plausible for a local-inference design, but the launch post offers no third-party audit. A practical check on macOS: run Foresight with a network monitor such as Little Snitch or the built-in Activity Monitor network tab and watch for outbound connections during a test meeting. Check whether it phones home for model downloads, analytics or update checks, and confirm which of those you can disable.
Recording consent. Capturing system audio is technically simple and legally uneven. Some jurisdictions require every participant to consent to recording. A bot-free recorder removes the visible signal that a call is being captured, so tell participants yourself.
Accuracy of enhanced notes. Expanding shorthand with retrieved context can hallucinate links between your fragment and the wrong part of the conversation. Treat the output as a draft and spot-check names, figures and commitments, exactly as you would with a cloud notetaker.
Hardware. Google's post says the app is a Mac download, and other coverage reports it is optimized for Apple Silicon. The post does not publish minimum memory or disk requirements, and we have not benchmarked it ourselves, so expect to find out by trying. If you already run a Gemma 4 model alongside other apps, watch memory pressure during long calls.
Experimental status. Google labels it experimental. Features can change or vanish, and there is no stated support commitment.
What developers can take from it
Even if you never install Foresight, the launch post doubles as a recipe. The ingredients Google lists are:
- LiteRT for the runtime, with a single
.litertlmfile running on CPU and GPU across supported edge platforms. - MediaPipe Tasks, which adds EmbeddingGemma 2 support to the Embedder and Semantic Retriever tasks and handles image resizing, tensor normalization and tokenization for you, returning 768-dimensional vectors that can be truncated to 128 to 512 dimensions.
- ML Kit on Android, coming in the next few weeks per Google, with NPU acceleration and automatic model updates.
- Pre-quantized bundles from the Hugging Face LiteRT Community, plus a Colab and GitHub guide.
A minimal local-search architecture, following Google's own description, is: embed every note, transcript chunk and image into one vector space, store vectors in a local database (the Gallery uses SQLite), and retrieve by cosine similarity as the user types. Add a small Gemma 4 model on top to rewrite or answer using the retrieved snippets. That is the same shape as Foresight, and the privacy story comes from never calling a hosted API.
A sensible way to trial it
Start with a low-stakes recording rather than a client call. Run a short internal meeting with Foresight capturing system audio, jot a few shorthand fragments as you go, and compare the enhanced notes against your own memory of what was said. Check three things: whether names and numbers survived intact, whether the expanded fragments point at the right moment in the conversation, and whether search over older notes and images returns what you expect when you type a loose query.
Keep the network monitor running throughout, and note memory pressure if a Gemma 4 model is also loaded. If the first two meetings hold up on accuracy and traffic, move to longer calls; if not, you have lost little. Pair the trial with one of the open alternatives above so you have a baseline to judge against.
Why this matters
For people who pay for cloud meeting assistants, a free local alternative from Google changes the default calculation, particularly for legal, medical, HR and finance conversations where uploading recordings is a compliance headache. For builders, it is more evidence that the useful part of a meeting assistant, retrieval over your own material, no longer needs a server. The unanswered questions are quality and trust: how good are Gemma 4-sized notes compared with frontier cloud models, and will Google keep the app maintained?
We will update this post once independent hands-on tests appear. If you try it, the checks worth publishing are an outbound traffic log, a side-by-side note quality comparison against a cloud tool, and memory use during a one-hour call.
Related reading
- EmbeddingGemma 2: open multimodal embeddings in 740M parameters
- Gemma 4 12B: multimodal local model for 16GB laptops
- Gemma 4 updates: flash attention and tool calling
- Meetily: privacy-first local meeting assistant
- FluidVoice: open-source macOS dictation
- Files.md: local-first note-taking
Official sources: Google developers blog on EmbeddingGemma 2 and AI Edge and Google AI Edge documentation.
Details reflect Google's announcement on October 8, 2026 and may change as the app updates.
