deniz.in

Markets

Weather

Loading weather

· via Hacker News – Front Page (native)

Open-weight security AI catalogue compares uncensored models from 2.6B to 180B parameters

A GitHub catalogue that hit Hacker News's front page compares 14-plus open-weight models fine-tuned or uncensored for offensive security, from a 2.6B classifier to a ~180B MoE build.

Open-weight security AI catalogue compares uncensored models from 2.6B to 180B parameters

A GitHub repository that rounds up open-weight language models fine-tuned — or deliberately de-restricted — for cybersecurity work has landed on Hacker News's front page, giving red teams and researchers a single, current comparison of a crowded niche. The list, maintained by developer JoasASantos and dated September 2026, tabulates more than a dozen models against their base architectures, parameter counts, context windows, quantised memory requirements, training data and licences.

How the catalogue is assembled

According to the repository, every figure comes from HuggingFace model cards and the publishers' own documentation, so the numbers are self-reported rather than independently verified. The stated scope is authorised use: penetration testing, red team exercises and security research. The range is wide. At the small end sits Cyber-Prime 1.1, a roughly 2.6-billion-parameter model built on LiquidAI's LFM2-2.6B that runs in about 6 GB of memory and reports a 0.592 average on its CyberBench suite, with phishing detection at 0.890. At the other extreme, Blackfrost-AI's CYBER-FROST-3.8 is listed at around 180 billion total parameters spread across 512 routed experts, with a 262K-token context window and a hardware appetite to match — the entry cites testing on four NVIDIA B300 accelerators.

Three routes to a security model

The catalogue shows three dominant adaptation strategies. The first is supervised fine-tuning on security corpora: WhiteRabbitNeo's DeepHat V2 is described as trained on 1.7 million offensive and defensive samples on top of Qwen2.5-Coder-7B; BugTraceAI-CORE-Apex draws on HackerOne disclosure reports from 2024 and 2025; and CyberPal 2.0, built on OpenAI's open-weight gpt-oss-20b, uses 403,000 examples assembled through an expert-in-the-loop pipeline.

The second is lightweight adaptation. Dolphin3-Cyber-8B applies a rank-16 LoRA over an already-abliterated Dolphin3 base, while pentest-v2 adds a rank-4 LoRA with roughly 2,800 curated samples drawn from GTFOBins, HackTricks and HackTheBox writeups — and claims perfect accuracy on GTFOBins questions versus 25 percent for the base model answering zero-shot.

The third is refusal removal itself. Cyber-Ornith-1.5-9B is an "obliterated" variant of an Ornith 9B release, aimed at agentic security work such as function calling, tool use and command-line automation, shipped as GGUF files from IQ1_S through Q6_K.

Specialisations split along offensive and defensive lines

Several entries target the offensive side. BugTraceAI-CORE-Ultra positions itself as a tooling model that generates Nuclei templates, CVE proof-of-concept code and pentest scripts from a comparatively small set of 2,541 training examples. RavenX-CyberAgent, a 36-billion-parameter mixture-of-experts build, follows a six-stage protocol spanning attack surface, exploitation, impact, remediation and reporting, and emits CVSS scores, CWE identifiers and MITRE ATT&CK mappings in its output.

Others lean defensive. CyberPal 2.0 covers threat intelligence, vulnerability analysis, detection, incident response and compliance, while Imperum-CybersecurityLLM spans SOC and SIEM operations, forensics, malware analysis, cloud, Kubernetes and identity security, OT security and governance. Qwythos-9B stands out for a one-million-token context window via YaRN rope scaling plus an inherited vision capability.

Licences and running costs

Most entries ship under Apache 2.0, but not all: CYBER-FROST-3.8 carries the Qwen Community License 1.0, Dolphin3-Cyber-8B the Llama 3.1 licence, and Cyber-Prime the LFM Open License v1.0 — a reminder that open weights do not always mean a permissive licence. Hardware demands scale accordingly. According to the listings, 7B-to-9B models need roughly 6–7 GB of VRAM at Q4_K_M quantisation, mid-size 26B–36B builds land between 16 GB and about 24 GB, and the largest entry requires multi-GPU servers.

Why it matters

Frontier chat models routinely refuse offensive-security queries, and professionals running authorised tests increasingly reach for open-weight alternatives. This catalogue documents what exists, at what size, trained on what data and under which licence — the practical questions that decide whether a model fits a given lab. Two caveats follow from its own methodology: the specs are self-reported by publishers, so benchmark claims such as pentest-v2's 100 percent GTFOBins score are leads to verify rather than settled results, and uncensored models carry real misuse risk if they escape authorised settings. As a snapshot of how quickly security fine-tunes track each new open base release — Qwen, Gemma, Llama, Mistral and gpt-oss all appear here — it is a useful bellwether for the field.

  • #open-source
  • #cybersecurity
  • #llm
  • #red-team
  • #hugging-face

Related posts