Security · Advanced

Why Monsmith has Qwen 3.8 Max cross-examine every finding before it reaches a report

How Monsmith uses Qwen 3.8 Max as an independent reviewer in its smart contract audit: why the reviewer comes from a different model family, the rule that it must cite a line to reject anything, what it did on The DAO, and what we had to fix to run it.

By , founder of Monsmith6 min read

A smart contract audit is trusted or it is ignored. One confident finding that turns out to be wrong costs more trust than three missed low-severity issues cost coverage, because it teaches the reader to skim. The hardest part of an AI audit is not finding things. It is not reporting the things that are not there.

Monsmith's answer is a second auditor from a different model family, whose only job is to attack the first one's findings. Since 11 October 2026 that second auditor is Qwen 3.8 Max.

Where it sits in the audit

Every contract goes through four passes, each feeding the next:

  1. Describe: what the contract does, in functional terms, with no security framing. Fast models.
  2. Threats: what an attacker would go after, derived from this contract's own assets, actors and invariants. Fast models.
  3. Hunt: concrete vulnerabilities, each tied to a line number and an attack that would work against this code. Claude Sonnet 5.5.
  4. Review: every candidate finding is judged by Qwen 3.8 Max, which may uphold it, correct its severity, or reject it.

Deterministic static analysis runs alongside. What it proves (for example, an external call made before a balance is zeroed) is handed to the reviewer as settled fact that it is not allowed to argue with.

Why a different family

Asking a model to check its own findings mostly returns agreement, because the reasoning that produced a mistake is the same reasoning being asked to catch it. Two models trained by different labs on different data make different mistakes. Putting Qwen behind Claude means a finding has to convince a reader that did not share the hunter's blind spots.

Qwen 3.8 Max suits the job for three practical reasons. It reads a whole contract with line numbers at once (its context is a million tokens). It reasons before it answers, which is what judging an attack path needs. And at about 2 dollars per million input tokens and 6 per million output tokens through OpenRouter, a review costs a fraction of a cent to a few cents per contract.

The rule that makes the review safe

Left to its own judgement, a reviewer will sometimes delete a real vulnerability with a confident-sounding sentence. So Monsmith's reviewer works under one hard rule: it may only reject a finding by citing the specific line that prevents the attack, such as a modifier, a require, an ordering or a type constraint. A rejection without a line number is not honoured, and the finding stands. Findings the reviewer does not mention are kept, and if the review fails to run, every finding is kept and the report credits no reviewer.

What it did on The DAO

The DAO is the start contract in Monsmith's studio: a rebuild of the 2016 contract whose splitDAO reentrancy drained about 3.6 million ETH. In a run on 11 October 2026, the review pass received three candidate findings and answered as follows.

Candidate findingQwen 3.8 MaxWhy
Drain ether during withdrawal (the splitDAO reentrancy)Upheld, criticalNo line prevents it: ether leaves before the caller's balance is burned.
Manipulate total supply to exceed the ether contributedRejectedCited line 24, require(msg.value > 0): the attack's second step, buying tokens with zero ether, cannot happen.
Change the curator to a malicious actorRejectedCited line 77: the attack assumes the curator's key is already compromised, which is a key-management problem, not a flaw in this contract.

The real, famous vulnerability survived. The two that did not survive were each rejected with a line a reader can check in seconds. That is the behaviour the reviewer is there for.

What the reader sees

The report's Health section says, for example: "Every AI finding was cross-examined by Qwen 3.8 Max: 4 findings upheld, 2 candidates rejected for lack of a working attack." The same line is in the downloaded Markdown. The model id is recorded from the API response that actually answered, not from configuration, so if Qwen is unavailable and another model stands in, the report names that model instead.

When a report is published to Monad, its keccak256 hash goes onchain with the score under Monsmith's ERC-8004 identity. The verification record is part of the hashed report, so who checked the findings, and what they decided, is covered by the attestation and cannot be edited afterwards without the hash ceasing to match.

Two things we had to fix to run it

  • Qwen 3.8 Max always reasons; the API refuses a request that turns reasoning off. With the default token cap it spent its budget thinking and timed out. It now runs with reasoning effort set to low and room for 16,000 tokens.
  • The same investigation found that Claude Sonnet 5.5 sometimes thinks before answering too. In one run it spent 2,341 of a 4,096 token cap reasoning, the findings JSON was cut off mid-sentence, and the audit parsed as empty. Every strong pass now has room, and a completion cut off at its cap fails over to the next model instead of being read as no findings.

Limits

A second model is a filter, not a proof. Qwen can uphold a finding that is wrong, and the citation rule means it will sometimes keep a weak one rather than risk deleting a real one. That trade is deliberate. For anything that holds real value, Monsmith's report is a strong first pass and a human audit should still follow.

Frequently asked questions

Which model does Monsmith use to review audit findings?
Qwen 3.8 Max, through OpenRouter. It cross-examines every finding produced by the hunting model, Claude Sonnet 5.5, and may only reject one by citing the line that prevents the attack.
Why use two different AI models in one audit?
A model checking its own work mostly agrees with itself. A reviewer from a different model family does not share the first model's blind spots, so false positives are caught without deleting real findings.
Does the report say which model checked it?
Yes. The report and its Markdown download name the model that actually answered the review and how many findings it upheld and rejected, and that record is covered by the report hash attested on Monad.

Try it in Monsmith

Free, in the browser, no account. Compile, audit, profile, and deploy to Monad.

Open the studio