RV32I pipeline: 10/10 runs on a low-cost model

Bring AI into IC design. Let the tools be the judge.

RTLoop helps chip teams adopt AI without trusting it blindly. In our verification loop, a low-cost model writes RTL one module at a time, and simulation, structural checks and synthesis decide what passes. Failures go back to the model as raw tool errors, within a fixed call and cost budget; if a module still fails, it stops and an engineer decides.

We start every engagement with a pilot on one real block.

run.stderr — rv32i-pipe-v2 · start A-01
Replay

Replay of a real run log: a low-cost model writes the 13 work packages of a 5-stage RV32I pipeline; failures are fed back until every package passes. 20 calls, 934.5 seconds.

Real log, abridged · ↳ lines from check logs This run: 13 packages · 20 calls · 934.5 s
10/10 5-stage RV32I pipeline runs passed, every module written by a low-cost model
0 → 10/10 same model and spec: one big prompt vs our module-by-module loop (RV32I single-cycle)
$0.0081 model API cost per passing run; excludes engineering time for specs and tests
97% of our reference RTL's Fmax on the RV32I pipeline, with 0.8% more area (pre-placement estimate)
Built on open, version-locked EDA
  • Icarus Verilog
  • Yosys
  • OpenROAD
  • OpenSTA
  • ORFS
  • KLayout
  • SkyWater SKY130

01How it works運作方式

A loop where the model can't grade its own work

Every work package runs the same cycle. The model only proposes; the verification tools decide. A passing revision is locked and becomes a read-only dependency for the next package. The loop is model-agnostic: the same gates judge every model, and a step beyond a low-cost model, such as a bus interface, can go to a stronger one under the same gates and budgets.

FIG 01 — The verification loop One work package
Engineer

Interface spec & tests

Ports, cycle-accurate tests and, where available, reference RTL that the tests must pass first.

design-check → PASS
Model

Propose RTL

Writes only its assigned modules and returns complete files in a single tool call.

submit_rtl_files()
Tools

Verify

Admission check, then compile, simulate and check structure: single driver, no latches.

PASSlock it
NEEDS_REPAIRsend errors back
STALLEDclean restart
BUDGETstop, ask an engineer
iverilog · vvp · yosys
Flow

Lock & continue

The passing revision is frozen. Dependent packages start, and integration packages assemble verified modules.

revisions/NNN

↺ NEEDS_REPAIR and STALLED go back to the model; when the budget is used up, the loop stops for an engineer.

When every package has passed, all checks run again on the exact final source before synthesis. The loop only guarantees that RTL passes your tests, so we test the tests too: first against reference RTL, then against seeded bugs (all 22 caught on the RV32I pipeline).

What the model sees

  • The interface spec of its module
  • The plan notes for its package
  • Verified dependencies, read-only
  • Its own previous version
  • Raw tool output from the last failure
  • The repair history of this package

What it can't do

  • Edit tests or other modules
  • Declare its own success
  • Exceed call or cost budgets
  • Use latches, initial blocks or system tasks

02Results實測結果

Measured, not promised

In every run in the table, one low-cost model wrote every module from an empty file. Our team prepared the interface specs, tests and decomposition plans beforehand; that preparation is where the expertise goes.

FIG 02 — Blank RTL to synthesis SkyWater SKY130HD
Design Modules Passed Cost / pass Time / pass Fmax Area
gcd88-bit GCD engine 2 8/8 $0.0002–0.0006 26–63 s 310–328 MHz 1,640–1,724 µm²
RV32I single-cycleinteger-subset CPU 7 10/10 $0.0063 ≈ 8 min 115.1 MHz 68,212 µm²
RV32I 5-stage pipelinesplit into stage modules 13 10/10 $0.0081 ≈ 11 min 177.9 MHz 85,030 µm²
SHA-256 coresplit plan with stall restarts 5 9/10* $0.0063 — — —
AES coreone architecture note added by an engineer (0/8 without it) 6 8/10 $0.0143 ≈ 12 min — —

Model-written RTL held up against our reference RTL. The split pipeline reached a median Fmax of 177.9 MHz against the reference's 183.3 MHz, with 85,030 vs 84,361 µm² of area. The single-cycle core beat its reference: 115.1 vs 105.2 MHz.

*SHA-256: latest round to synthesis (2026-09-28). Without stall restarts, the same plan scored 6/10, 10/10 and 6/10 in three rounds. RV32I is the base integer subset (FENCE, ECALL and CSR treated as no-ops), checked with our own testbenches rather than the official riscv-tests; memories sit outside the core. Fmax and area come from ORFS synthesis and OpenSTA before placement, with Fmax = 1000 / (clock period − worst setup slack); medians where several runs exist. Engineering estimates from September–October 2026, not signoff and not a formal model qualification.

Scope today: synthesizable Verilog-2005 with a single clock, simulated with Icarus Verilog and synthesized with Yosys and OpenROAD on SkyWater SKY130.

Why expertise matters

Decomposition decides success

Same pipeline spec, same low-cost model. With the whole 5-stage RV32I core in one package, most runs failed. Split into stage modules with cycle-by-cycle tests, every run passed, at about a ninth of the cost.

What we changed

  1. One small module per package: tens of lines, not hundreds.
  2. Large sequential blocks become one stage of logic plus its own registers, each with its own test.
  3. Registers are declared on the module interface and cleared on reset, so every cycle can be compared.
  4. Tests stop at the first mismatch and print every input, actual and expected value.
  5. Interface specs are exact to the cycle: who updates when, reset priority, edge cases.
FIG 03 — Same pipeline, two decompositions 10 runs each · Fisher p = 0.003
Whole core in one package Split into stage modules
Runs passedhigher is better
Whole core in one package: 3 of 10 runs passed
Split into stage modules: 10 of 10 runs passed
Model cost per passlower is better
Whole core in one package: 0.074 US dollars
Split into stage modules: 0.0081 US dollars
Median time per passlower is better
Whole core in one package: 30 minutes
Split into stage modules: 11 minutes
32k output-token overrunslower is better
Whole core in one package: 51 overruns
Split into stage modules: no overruns

Physical design

Beyond synthesis: RTL to GDS

We have also taken three blocks with prepared physical profiles to GDS through a fixed, hash-locked SKY130 flow. These were earlier, human-led runs: SPI's RTL is a low-cost model's refactor of reference RTL. For AES-128 and SHA-256, a low-cost model wrote the arithmetic blocks and stronger models wrote the rest, including the AES core, both APB interfaces and SHA-256's compression, padding and stream blocks. RV32I and the other designs above stop at synthesis today.

FIG 04 — Physical flow 17 acceptance conditions
  1. synthmap + equivalence
  2. floorplancore & pins
  3. placecell placement
  4. ctsclock tree
  5. routesignal wiring
  6. finishGDS out
  7. DRC0 violations
  8. LVSMATCH
Metal layers of the SHA-256 / APB digital core, read back from the delivered GDS
SHA-256 / APB digital core: metal layers read back from the delivered GDS. 957 × 957 µm, 100 ns clock, 17/17 conditions (2026-09-09).

Taken to GDS

SPI17/17 AES-128 / APB17/17 SHA-256 / APB17/17

AES-128 and SHA-256 at a 100 ns clock (10 MHz). Human-led runs, several models.

Required to pass

  • setup slack ≥ 0
  • hold slack ≥ 0
  • DRC = 0
  • LVS = MATCH

Plus 13 more conditions, 17 in all.

Digital-core layouts on the open SkyWater SKY130 PDK. They show the flow runs end to end; pad ring, packaging and tape-out signoff are out of scope.

03Services服務

Two ways to bring AI into your flow

Whether your own engineers will run AI-assisted design or you want a block delivered, we work the same way: measurable gates first, models second.

Consulting

AI adoption for IC design

We review your design and verification flow, find the routine RTL work a model can safely take over, and build the gates, budgets and audit trail around it.


What we cover

  • Flow assessment and adoption roadmap
  • Specs and testbenches written for AI
  • Model selection and cost control
  • IP-aware deployment
  • Hands-on training for your engineers
Projects

Turnkey blocks, delivered with evidence

Give us a block and its spec. We write the interface specs and tests, plan the decomposition, run it through the loop and deliver RTL that has passed every tool check.


What you receive

  • Design kit: interface specs, testbenches and a mutation-test report
  • Decomposition plan and the accepted RTL of each module, with digests
  • An HTML report of every model call and its cost
  • Synthesis results: timing, area, power, netlist and SDC

How an engagement runs

Each step ends with numbers you can check.

  1. 01

    Discover

    Understand your flow, tools, IP constraints and where engineering time actually goes.

  2. 02

    Pilot

    Run one real block through the loop and measure pass rate, cost and time.

  3. 03

    Integrate

    Connect the loop to your regressions, reviews and budgets.

  4. 04

    Hand over

    Playbooks and training so your team can run it on its own.

04IP & security智慧財產與資安

You control what leaves

Chip teams can't afford a leaked design. Only the minimum context for one module leaves your machines, and every request is logged.

01

Verification stays local

Simulation, structural checks, synthesis and the physical flow run on your machines. Reference RTL, testbenches, netlists and layouts never leave.

02

Minimum context per call

A request carries one package's interface spec and plan notes, the RTL of its verified dependencies, its previous version, and tool feedback with the repair history. Nothing beyond that package.

03

One pinned provider

Today, model calls go through OpenRouter to one pinned inference provider, with fallback routing off and data collection denied.

04

Every call on record

Each request and response is stored with the project for audit. API keys live only in the subprocess that calls the model, never in logs, reports or Git.

05Team團隊

IC design and AI, in one team

RTLoop was founded in Taiwan on June 5, 2026. Our 10 members include graduate students in IC design and AI from the Department of Electrical Engineering and the Department of Computer Science and Information Engineering at National Taiwan University of Science and Technology (NTUST). We track AI-for-EDA research, test new ideas on our own benchmarks before relying on them, and hold our own work to the rule we bring to clients: if a tool didn't verify it, it doesn't count.

  • 10members
  • NTUSTEE · CSIE
  • 2026founded in Taiwan

How we got here

  1. RTLoop founded
  2. Development repository started: 813 commits and 68 dated result reports by October 2
  3. AES-128 / APB taken to GDS, 17/17 (human-led)
  4. SHA-256 / APB taken to GDS, 17/17 (human-led)
  5. RV32I single-cycle from blank RTL: 10/10 on a low-cost model
  6. RV32I 5-stage pipeline: 3/10 → 10/10 after splitting into stage modules
Results and evidence on GitHub

06Contact聯絡

Start with one block.

Tell us about your design flow. We'll propose a pilot with agreed targets for pass rate, cost and turnaround, and run it on a real block.

boss@rtloop.com