Bitcoin’s AI security sprint found 6,700 issues in 55 hours, but no one knows how many are real

AI-assisted security campaign focused on the Bitcoin ecosystem, Bitcoin Red Team, said it generated 6,700 findings across 425 projects in its first 55 hours. The campaign labeled 1,029 of them high or critical.

The Aug. 6 update measures how much material entered a security triage pipeline, and its effect on software security remains unreported.

The retrieved thread omitted audit-ready definitions and denominators for the severity counts, as well as case-level outcomes, an aggregate false-positive rate, and a fix rate.

Those missing fields prevent a calculation of how many alerts became confirmed vulnerabilities, how many maintainers rejected or downgraded, and how many led to patches.

The first 55 hours still reveal a consequential capability, noting how AI systems can fill an ecosystem-scale review pipeline quickly. Expert prompting, reproduction, disclosure, and maintainer response remained necessary at every later stage.

What the campaign numbers measure

The campaign published two snapshots as its roster and workload expanded:

Elapsed time Projects Total findings Reported severity Participants
27.5 hours 390 4,962 85 critical; 635 high 16
55 hours 425 6,700 1,029 high or critical 24 reported, including three bots

The 27.5-hour update covered 390 projects and 4,962 findings. By the 55-hour mark, the project count had risen by 35 and the finding count by 1,738. The later thread put high-or-critical findings at 15.4% of the total and clarified that three of the 24 reported participants were bots.

Read More:  Ethereum Foundation cuts 20% of staff as ETH sinks 44% YTD despite record usage

The earlier post separated critical and high findings, while the later one combined them, with both sets of figures reflecting campaign assessments. Maintainer-confirmed exploitability and remediation outcomes require separate evidence.

Bitcoin Red Team scanned 425 projects and reported 6,700 findings, including 1,029 high or critical issues, while public validation rates remain unpublished.

Rob Hamilton described Kimi K3 as handling the heavy analysis, with GPT Sol, Fable/Opus, and GLM 5.2 supporting the documentation. He said OpenAI’s Cyber Harness covered selected components he considered load-bearing.

A day later, Hamilton wrote that subject-matter experts could change an assessment with one or two sentences of context or a small block of code. In examples he described, that input pushed middling concerns into high or critical territory. He also identified operations, disclosure handoff, and triage as bottlenecks.

In Hamilton’s account, models searched broadly while specialists shaped prompts, interpreted output, attempted reproduction, and decided which reports were ready for disclosure. That division of labor makes the campaign a human-AI review system.

Related Reading

OpenAI’s new cybersecurity push has a lesson for crypto: stop waiting for the hack

OpenAI’s Daybreak may point to the crypto industry’s next security standard of becoming resilient before vulnerabilities are exploited.

May 12, 2026 · Gino Matos

The developer known as Calle said most critical reports were quickly verified by project owners. The post supplied no denominator, verified-report count, rejection count, or patch status, leaving the breadth and outcome of that verification unresolved.