deniz.in

Markets

Weather

Loading weather

· via Hacker News – Front Page (hnrss.org)

Screen Memories brings on-device AI search to every photo and video frame on macOS

SCM, an open-source macOS app that reached Hacker News's front page, indexes photos and every video frame locally, offering natural-language, OCR and spoken-dialogue search that never leaves the machine.

Screen Memories brings on-device AI search to every photo and video frame on macOS

What shipped

A developer has released SCM, short for Screen Memories, an open-source macOS application that runs AI models over personal media libraries so users can find photos and individual video frames using plain-language descriptions. The project reached Hacker News's front page on October 4 via a Show HN post pointing to its GitHub repository, and its README lays out the central promise: model inference happens on the machine itself, with no account required and no media uploaded anywhere. Weights are fetched once at first use; everything afterward is offline.

The app is an Electron build distributed through a Homebrew tap for Apple Silicon Macs running macOS 12 or later, with version 0.2.4 currently shipping as DMG and ZIP installers. Development uses Bun as the package manager and runner, and the initial vision-model download is roughly 435MB.

Five ways to search

According to the GitHub README, retrieval is split into five modes:

  • Files — semantic ranking over whole photos and videos. A filename-keyword pass returns instant results, then the vision model scores images by cosine similarity against embeddings, with per-model calibration to keep similarity scores meaningful and a filter that suppresses near-duplicates. Each result tile explains which signal drove the match, and CJK queries are processed as overlapping bigrams.
  • Scenes — ffmpeg detects shot boundaries and builds a segment plan for each video; every segment's midpoint frame is embedded. A hit lands on the exact shot with its timecode, and opening the result seeks straight to that moment. A noise gate returns no match rather than flooding results, and each video contributes at most three scenes.
  • OCR — literal matching of query tokens against text extracted by Tesseract, with English plus 35 optional languages. Since no vision model is involved, it keeps working while the AI engine warms up or is unavailable.
  • Dialogue — exact spoken-word retrieval over Whisper transcripts at three levels of precision, from a contiguous phrase within one utterance down to all words appearing somewhere in the same video. It needs no embeddings and survives engine downtime.
  • LLMs — a strictly opt-in chat layer. A llama.cpp sidecar bound to loopback answers questions using evidence the app has already extracted, with numbered, clickable citations and streamed output. Qwen3 1.7B (1.1GB) is the default model; Llama 3.2 3B (2GB) is offered for longer answers.

Video indexing density is adjustable, from one embedding point every 60 seconds in the Eco preset down to one every 2.5 seconds in Ultra Pro, with each preset showing its measured time and disk cost before committing.

Models and library maintenance

Four vision models run through ONNX Runtime and are switchable per library. CLIP ViT-L/14@336 is the default at roughly 480–570ms per image on CPU, while the SigLIP variants trade between speed and detail: SigLIP-2-B/16 is fastest at 50–100ms per image for bulk imports, and SigLIP-B/16@384 targets maximum detail. Switching models re-embeds the library in the background while search temporarily falls back to filename keywords.

The library largely maintains itself, per the README: watched folders import automatically, content hashes deduplicate files even after renames, and any query can be pinned as a saved tab alongside built-in views for videos, screenshots and email. The Email tab surfaces photos whose visible text contains an address, reassembling ones that Tesseract fractures across word boxes and tolerating obfuscations such as allen [at] gmail [dot] com. The Screenshots tab classifies rename-proof using a priority chain of manual override, localized screenshot filenames, PNG and EXIF metadata probes, and folder hints.

Why it matters

Much of the AI photo search consumers encounter today is tied to cloud platforms, which means personal media must be uploaded or analyzed under someone else's control. SCM is a working demonstration that vision-language models, speech transcription and small local LLMs can now be combined into a full-frame, full-text search experience entirely on a consumer laptop, complete with citations and timecode-level video jumps. The costs are tangible — hundreds of megabytes of weights, CPU-bound embedding times and a self-maintained index — but for anyone with large local archives or privacy-sensitive collections, it moves a previously cloud-gated capability onto the device itself, a pattern likely to spread as on-device models keep improving.

  • #macos
  • #local-first
  • #image-search
  • #open-source
  • #machine-learning

Related posts