Install AgentBook a Call
All posts

August 8, 2026 · 2 min read

DeepSeek V4 Flash 0731 Hits 89% on ARC-AGI-1 at $0.02 Per Task

DeepSeekAI BenchmarksBusiness Automation

DeepSeek has published ARC Prize benchmark results for DeepSeek V4 Flash 0731, detailing its performance and cost-efficiency across multiple reasoning tiers.

On the ARC-AGI-1 Semi-Private evaluation, DeepSeek V4 Flash 0731 scored 89.0% at a cost of just $0.02 per task. On the ARC-AGI-2 Semi-Private track, it achieved 61.4% at $0.04 per task. Both of these benchmarks reflect the model operating at its maximum effort tier.

The Performance Breakdown

The published data breaks down the model’s performance across three distinct reasoning variants: Max, High, and Low.

The ARC-AGI-1 public evaluation tested the model against 400 tasks. Here, the performance degradation between reasoning tiers was relatively narrow:

  • Max variant: 89.0%
  • High variant: 87.0%
  • Low variant: 84.0%

The ARC-AGI-2 public evaluation, which included 120 tasks, showed a much steeper drop-off when compute effort was reduced:

  • Max variant: 61.4%
  • High variant: 56.0%
  • Low variant: 46.0%

This granular pass/fail data per reasoning level provides a clear picture of how the model scales its capabilities based on the resources allocated to a task.

What This Means for SMB Automation

DeepSeek is a major frontier player, and its model releases directly impact the cost-efficiency calculations for AI-automation partners. For small and mid-sized businesses looking to streamline operations, the numbers attached to DeepSeek V4 Flash 0731—specifically the $0.02 to $0.04 per-task cost at max effort—are highly relevant.

When our team builds automations for admin and operational bottlenecks, unit economics are the deciding factor. A "task" in a benchmark translates to a concrete business action in the real world: routing a complex customer inquiry, reconciling line items on a vendor invoice, or structuring messy CRM data.

If an AI model can apply its highest tier of reasoning to a task for four cents, automation becomes viable for processes that were previously too expensive or too complex to hand off to software.

Furthermore, the three reasoning variants (Max, High, Low) allow for highly optimized workflow design. Not every administrative task requires maximum compute. A basic data entry or formatting step might run flawlessly on the "Low" variant, costing fractions of a penny. Later in the same workflow, a step requiring complex logic or exception handling can automatically switch to the "Max" variant to ensure accuracy.

The release of DeepSeek V4 Flash 0731 proves that high-level model performance is becoming increasingly cost-effective. For SMBs, this means the barrier to automating complex, multi-step administrative operations continues to fall.

Inspired by this source.

Stop doing the busywork yourself

A dedicated AI technician — backed by a full engineering team — automates your admin inside the tools you already use.