01 / Computer use and control
Compass
See the screen. Choose the move. Check what changed.
Compact models for understanding a screen, choosing the next action, and knowing when to hand off or ask for approval.
- 01Observe
- 02Choose
- 03Verify
Zenix AI / research & models
One family reads the screen. One maps the code. One shapes the interface. We are building each around a different kind of judgment.
03
model families
14
named variants
01
research snapshot
Three families / three kinds of work
These lineups describe Zenix's direction. Planned sizes and capabilities are not release dates or availability promises.
01 / Computer use and control
See the screen. Choose the move. Check what changed.
Compact models for understanding a screen, choosing the next action, and knowing when to hand off or ask for approval.
02 / Coding and software engineering
Map the change before it lands.
Models planned for code understanding, software engineering tasks, and increasingly capable agentic development workflows.
03 / Taste, design, and frontend work
Treat every interface as a whole.
A planned family focused on visual judgment, interface design, and building polished frontend experiences.
Parameter counts do not indicate memory requirements. Roadmap content may change.
02 / Research snapshot · CALIB1
After refitting temperatures on a validation split, Compass-1 v3 improves on some hard sets. It also regresses on a clean set and remains overconfident outside its training distribution, so it was not promoted.
v3 before refit
0.450
v3 after refit
0.022
Shipped model
0.033
Julia7
0.024
Expected calibration error on held-out in-distribution rows; v3 refit temperatures were fitted on a separate validation split. Lower is better.
CALIB1 results after validation-split temperature refit · scale 0–100%
VAULT1 hard
DESK1 hard · desktop
Keyboard slice
test47
orig170
jev791
newtest
Internal experiment results supplied by the Zenix team; not an independent audit or a cross-provider leaderboard. Scores compare Compass-1 v3 with the shipped Compass-1 on the named sets. Confidence intervals and significance are shown only where the run summary provided them.
The original release rule remains unchanged. Better in-distribution calibration and hard-set accuracy do not offset the current out-of-distribution and clean-set issues.
Calibration outside distribution
ECE is 0.221 on test47 and 0.116 on the DESK1 desktop hard set, both above the 0.10 bar.
Desktop abstention
Desktop none-on-real is 12.3% against a 2% bar. This is a major improvement over the shipped model's 91.8%, but still misses the release rule.
Clean-set regression
On newtest, v3 scores 94.40% versus 97.80% shipped. The −3.40-point change is significant (CI [−5.4, −1.6]).
Escalation and quantization
v3 escalates only 0.4% of rows, versus 10–23% for shipped on the same sets. Its int8 export also loses up to 6.4 points on test47 and 5.9 on orig170.
03 / Systems
Ora combines computer-use models with screen understanding, action checks, and approval boundaries. ZenixDash explores opt-in AI diagnostics for hosting operations.