deniz.in

Markets

Weather

Loading weather

· via Hacker News – Front Page (native)

Developer says Meta's Muse assistant is a cheap, effective tool for large-scale web scraping

A developer building a concert archive says Meta's Muse assistant can crawl Reddit and YouTube for thousands of videos a week on an $80 plan, raising questions about AI agents scraping at scale.

Developer says Meta's Muse assistant is a cheap, effective tool for large-scale web scraping

A developer has found an unexpected use for Meta's Muse assistant: large-scale web scraping. In a post on sigh.dev that surfaced on Hacker News' front page, the author describes pointing Muse at Reddit and YouTube to harvest thousands of concert recordings for a side project, consuming a few billion tokens a week on an $80 monthly plan — and often finding that the usage is not metered at all.

An assistant repurposed as a crawler

The author has little patience for Muse's pitch as a personal assistant. Ordering food or managing a calendar does not apply to their life; what they actually need is data, plus large amounts of inexpensive AI processing. That data feeds fullsets.fm, a project that gathers professionally recorded live performances — festival livestreams, Tiny Desk sessions and the like — organises them by artist and extracts tracklists.

The hard part was never judging videos. The author already runs a pipeline with models called Luna and Gemini 3.8, accessed via Gemini Studio, that cheaply decides whether a video has the quality and content they want. The gap was discovery: good recordings sit scattered across obscure YouTube channels — the post cites a T-Mobile Europe account that uploaded a handful of random concerts as an example — and there was no good way to find them all.

What the $80 a month buys

According to the post, Muse plugs both ends of that problem. It can be told to scan Reddit and YouTube for 1,000 videos matching the author's criteria, or to build a list of every artist who has played Coachella and then work through the list one artist at a time, searching YouTube for full concerts. It then sits at the far end of the pipeline, matching artists, reviewing what the other models produced, and making the final call to approve or reject each video.

The jobs run in the cloud on a managed virtual machine with Chrome, backed by what the author describes as generous bandwidth and a decent CPU, until they finish. The economics are the striking part: a few billion tokens a week on the $80 plan, with usage frequently appearing to go uncounted. The underlying model, Muse Spark 1.3, is described as poor at writing and unusable for coding, but well matched to this high-volume, low-stakes workload — the sort of thing usually delegated to cheap models like Gemini Flash.

Rough edges, and a warning

The post is candid about the friction. Jobs sometimes freeze partway through, communicating with the model can be frustrating, and the interface stops making sense when several browser-based jobs run at the same time.

The more consequential point is defensive. The author worries about what happens when someone aims Muse at their own websites: these agents never sleep, can convincingly pass as real users, and can be scaled up by launching more processes whenever a single one is too slow. The author notes that Amazon has already blocked Muse and expects more sites to follow suit unless Meta clamps down on abuse.

Why it matters

Muse was framed as a consumer assistant, but the demand it is actually meeting here is cheap agentic compute: a browser, bandwidth and effectively unmetered tokens in one package. That combination gives a hobby project industrial crawling capacity, and it previews a broader standoff between website operators and AI agents that can imitate human traffic at scale, complete with bot-block lists targeting a single vendor's agent. It also leaves Meta with a dilemma: the unmetered usage that makes Muse genuinely useful for this kind of work is precisely what invites the abuse and blocking that could end it.

  • #ai-agents
  • #web-scraping
  • #meta
  • #llm

Related posts