· via dev.to (home feed)
MongoDB Community search index stuck pending above 89% disk use in self-hosted tests
A dev.to write-up of the newly GA MongoDB Community search stack found indexes stuck in PENDING past roughly 90% disk use, with df and mongot disagreeing on capacity by 60 points.

What happened
MongoDB Search and Vector Search for Community Edition reached general availability on 30 June 2026, bringing the mongot search process, previously an Atlas-only component, to self-managed deployments where it runs alongside mongod rather than on a separate engine. A detailed write-up on dev.to describes following MongoDB's Docker installation guide to the letter: a mongod container (Community Server 9.0.2) and a mongot-community container (mongot 1.70.5) connected over gRPC with a SCRAM user holding the searchCoordinator role, a single-member replica set, and createSearchIndex against a 50,000-document test collection. The index came back PENDING and never left.
According to the post, a brief PENDING phase is expected. What was not expected was the same status five minutes later, with mongot logging transient sync errors and retrying every 30 seconds while initial sync stayed paused.
The disk threshold that stalls the build
The author found the cause in mongot's Prometheus metrics, exposed on port 9946: 28.7 GB free of 270.6 GB, or 89.4% used. MongoDB's own self-managed troubleshooting page states that replication halts once disk use passes roughly 90% and only resumes below roughly 85%, and that an index definition will be accepted but its build held up when disk pressure already exceeds that protective line.
Diagnosis was made harder by a measurement mismatch. On the same path, df -h reported 28% used, about 60 points lower, because it measured within a quota-limited allocation while mongot divided by the full capacity the underlying device reports. Nothing about the block surfaced client-side: createSearchIndex returned normally and $listSearchIndexes simply reported PENDING.
After the author moved mongot's data directory onto tmpfs to get a mount with headroom, the same index finished its initial sync in 4.85 seconds and transitioned to steady state, confirming that disk headroom alone was the blocker.
Status fields and silent failure modes
The post documents two reporting problems operators should know about. After the mongot container was recreated, $listSearchIndexes kept reporting the index as PENDING even though queries against it worked. The aggregated top-level status reflects the worst state across every host ever recorded, including ones that no longer exist, so the per-host statusDetail array is the trustworthy view.
Failure behavior is also asymmetric. Querying an index that exists but is not yet ready raises a hard OperationFailure with code 8, noting a NOT_STARTED state. Querying an index name that does not exist at all returns zero results silently, with only a warning line in mongot's own log, meaning a typo'd index name and a genuinely empty result set look identical from the application side.
Performance and cost findings
Once the index was live, the author ran 15 repetitions each of single-word and two-word queries on the same collection. $search through mongot posted a 9.36ms median, against 0.67ms for a classic $text index and 0.79ms for a case-insensitive regex scan, roughly 10 to 14 times slower for plain term lookups, which the post attributes to the gRPC round trip. Fuzzy matching is where the separate process pays off: a query for "backpak", a one-letter typo, returned five results with maxEdits set to 1, while $text, regex and non-fuzzy $search all returned zero.
For vectors, $vectorSearch on a 32-dimension cosine-similarity index ran a 15.52ms median against 134.89ms for a brute-force Python scan of all embeddings, about 8.7 times faster, and the comparison was generous to brute force because the wire fetch of all 50,000 vectors was excluded.
Concurrency improved throughput but hurt tails: ten clients completed 20 queries in 98.6ms of wall time versus 180.7ms sequentially, while median latency rose from 8.41ms to 28.77ms and the worst case from 13.07ms to 83.40ms. At idle, mongot held 1.033GiB resident against mongod's 229.5MiB, a second process with a JVM-scale footprint that does not shrink for small collections. Chaining $match, $project, $sort, searchScore metadata and $limit after $search worked without special handling. The post also records a dead end: suspecting the documented secondaryPreferred read preference on a single-node set, the author added a secondary replica set member, and the index still did not move.
Why it matters
Self-hosted adopters of the newly GA stack need monitoring aimed at mongot, not just mongod. The roughly 90% disk ceiling is invisible from the client, and on quota-backed mounts standard tools like df can disagree with mongot's own arithmetic by tens of points, so a capacity problem presents exactly like a broken feature. Capacity planning matters too, since search now carries a gigabyte-scale companion process that sits idle at over four times mongod's footprint. Reading per-host statusDetail rather than the summary status, and watching mongot's Prometheus metrics, are the practical takeaways before rolling the feature out.
- #mongodb
- #search-index
- #self-hosted
- #vector-search
- #database