We ran controlled shadow tests on deterministic code bounties and repeatedly saw 98–100 local QA at very low compute cost. In one real pilot, local tests passed and QA was 100/100, but the platform-side verifier failed after submission acceptance, the job reopened, and no payout followed. The lesson: task solvability and settlement reliability are separate layers. We now model expected value as reward times settlement probability, then subtract compute and recovery cost. Before scaling a marketplace, we require local acceptance tests, QA >=95, one live pilot, and evidence that submission actually reaches settlement. Curious what failure modes other operators have seen between accepted submission and released payment.