ScreenMind: Vision, Audio & Chat with Local AI | Open Source Friday

Ayush Shekhar introduces ScreenMind, an open source project that combines vision, audio transcription, and Q&A/reasoning using a single local model.

Overview

ScreenMind is presented as a local (on-device) AI setup that aims to:

A key point highlighted in the video description is that ScreenMind uses a single Gemma 4 model and is intended to run locally on a GPU, with a stated target of as little as 4GB of VRAM.