Zenix

Zenix AI / research & models

Intelligence for the work it touches.

One family reads the screen. One maps the code. One shapes the interface. We are building each around a different kind of judgment.

03

model families

14

named variants

01

research snapshot

A compass rose, needle, and cardinal marks formed from muted pastel dots

Three families / three kinds of work

Different questions need different instruments.

These lineups describe Zenix's direction. Planned sizes and capabilities are not release dates or availability promises.

A compass rose, needle, and cardinal marks formed from muted pastel dots

01 / Computer use and control

Compass

See the screen. Choose the move. Check what changed.

Compact models for understanding a screen, choosing the next action, and knowing when to hand off or ask for approval.

  1. 01Observe
  2. 02Choose
  3. 03Verify
Explore Compass 5 named variants
A dotted globe with route markers, curved meridians, and a stand

02 / Coding and software engineering

Atlas

Map the change before it lands.

Models planned for code understanding, software engineering tasks, and increasingly capable agentic development workflows.

  1. 01Trace
  2. 02Reason
  3. 03Patch
Explore Atlas 4 named variants
Four muted pastel dot-map tiles arranged as a diamond

03 / Taste, design, and frontend work

Mosaic

Treat every interface as a whole.

A planned family focused on visual judgment, interface design, and building polished frontend experiences.

  1. 01Compose
  2. 02Systemize
  3. 03Refine
Explore Mosaic 5 named variants

Parameter counts do not indicate memory requirements. Roadmap content may change.

02 / Research snapshot · CALIB1

A calibration fix is not a release verdict.

After refitting temperatures on a validation split, Compass-1 v3 improves on some hard sets. It also regresses on a clean set and remains overconfident outside its training distribution, so it was not promoted.

v3 not shipped

v3 before refit

0.450

v3 after refit

0.022

Shipped model

0.033

Julia7

0.024

Expected calibration error on held-out in-distribution rows; v3 refit temperatures were fitted on a separate validation split. Lower is better.

Accuracy by evaluation set

CALIB1 results after validation-split temperature refit · scale 0–100%

Compass-1 v3Shipped Compass-1

VAULT1 hard

0%50%100%
92.37%83.94%+8.43 pp

DESK1 hard · desktop

0%50%100%
76.07%70.05%+6.02 pp

Keyboard slice

0%50%100%
65.7%44.4%Not reported

test47

0%50%100%
89.36%91.49%−2.13 pp · CI [−6.4, 0.0] · not significant

orig170

0%50%100%
91.18%88.82%+2.35 pp

jev791

0%50%100%
86.85%88.87%−2.02 pp · CI [−4.4, +0.5] · not significant

newtest

0%50%100%
94.4%97.8%−3.40 pp · significant regression

Internal experiment results supplied by the Zenix team; not an independent audit or a cross-provider leaderboard. Scores compare Compass-1 v3 with the shipped Compass-1 on the named sets. Confidence intervals and significance are shown only where the run summary provided them.

Promotion review

Why v3 stays in research.

The original release rule remains unchanged. Better in-distribution calibration and hard-set accuracy do not offset the current out-of-distribution and clean-set issues.

Calibration outside distribution

ECE is 0.221 on test47 and 0.116 on the DESK1 desktop hard set, both above the 0.10 bar.

Desktop abstention

Desktop none-on-real is 12.3% against a 2% bar. This is a major improvement over the shipped model's 91.8%, but still misses the release rule.

Clean-set regression

On newtest, v3 scores 94.40% versus 97.80% shipped. The −3.40-point change is significant (CI [−5.4, −1.6]).

Escalation and quantization

v3 escalates only 0.4% of rows, versus 10–23% for shipped on the same sets. Its int8 export also loses up to 6.4 points on test47 and 5.9 on orig170.

03 / Systems

Models need a capable, careful operator.

Ora combines computer-use models with screen understanding, action checks, and approval boundaries. ZenixDash explores opt-in AI diagnostics for hosting operations.