Coming soon · in training

Ysra SWE

Our first software-engineering model — trained on how work actually gets finished.

Frontier models already know how to code. Ysra SWE is being trained on something rarer: real engineering work that was run, checked and proven — so it learns to finish the job, not just write the code.

Open-weightfoundationVerified Ysraengineering workYsraSWE
  1. Foundation selected and verifiedDone
  2. Training pipeline validated end to endDone
  3. Data pipeline and verifiers builtDone
  4. Training and evaluationUnderway
  5. Reinforcement learning on real tasksNext
  6. ReleaseComing soon

Why build our own model

Knowing how to code isn’t the hard part. Finishing is.

A general modelYsra SWE
Writes code that looks rightRuns it, tests it, and shows what happened
Says “done” when the answer reads wellSays “done” when the checks agree — and names what it couldn’t prove
Learns from text on the internetLearns from engineering work that was verified to hold up

The data behind it

We keep only what holds up.

Ysra has already done a great deal of real engineering. Very little of it is imitated as-is — only work that was checked and proven. The rest becomes tests, lessons and harder tasks.

  1. EverythingEngineering records captured

    Sessions, outcomes and traces from real Ysra work

  2. CompleteComplete, usable runs

    Whole sessions with a known outcome

  3. ComparedBetter-vs-worse pairs

    Two attempts at one task, one clearly better

  4. MinedFailures mined into lessons

    Where “done” turned out not to be

  5. RebuiltReproducible task environments

    Real repository, exact starting commit, hidden checks

  6. ImitatedVerified demonstrations

    Work that passed every check — the only runs we imitate

Held out

A set of real tasks, kept apart. Never trained on, never tuned against — kept aside so the release score means something.

Illustrative, not to scale. Each stage keeps less than the one before it.

How work is scored

Green tests aren’t the finish line.

A model trained to make tests pass will learn to make tests pass — even by deleting them. So Ysra SWE is rewarded by checks it can’t see, and gets nothing for a shortcut.

The kind of check we use

Task: rename calc_total to compute_total, keeping the old name working.

billing/totals.py
def compute_total(items): …

def calc_total(items):
    return compute_total(items)
  • Visible tests pass
  • Hidden check fails: the old name must be the new function. A wrapper isn’t an alias.
  • Reward0

What earns reward

  1. 01
    Hidden task checksThe behaviour you asked for, tested where the model can’t see.
  2. 02
    The repository’s own checksTests, types, lint and build — as your team runs them.
  3. 03
    Shortcut detectionDeleted tests, removed assertions, disabled checks: zero reward.
  4. 04
    Task-specific behaviourDoes it do the thing, not just pass the thing.
  5. 05
    Clean changeA small bonus for focused diffs — never enough to rescue broken work.

Break the core behaviour and nothing else counts. A neat diff can’t rescue broken work.

How it learns

Every verified lesson feeds the next model.

First it learns from work that already held up. Then it practises on fresh, real tasks — and only attempts that pass hidden checks shape the next version.

  1. 01FoundationA strong open-weight modelDone
  2. 02Learn from verified workImitate only runs that held upUnderway
  3. 03Attempt fresh tasksReal repos, exact starting commitsNext
  4. 04Score with hidden checksReward only what really worksNext
  5. 05Reinforce what workedUpdate, re-measure, repeatNext

What training is aiming for

Move the whole curve — not just one point.

More tasks solved, for less. Whether you run it lean or let it dig deep, Ysra SWE is built to get more real engineering done at every level of effort.

  • EconomyQuick, well-scoped work at the lowest cost.
  • BalancedThe everyday default for real engineering tasks.
  • IntensiveMore room to investigate and repair hard problems.
Starting modelYsra SWE · target
Illustrative target · not measured
Training aims to move the cost and tasks-solved frontier up and to the leftIllustrative chart with no measured values. The starting model's curve is shown dashed; Ysra SWE's target curve sits higher and further left, meaning more tasks solved at lower cost for each compute profile: Economy, Balanced and Intensive.Cost per task →Tasks solved →↖ better: more solved, lower costEconomyBalancedIntensive
An illustration of what training is designed to do — not a result. Real measurements publish with the release.

How we’ll judge it

It ships when it earns it.

Ysra SWE will be measured against its own starting point on a held-out set of real engineering tasks it has never seen, through the same Ysra agent you use.

  • It must solve more tasks.
  • It must not claim success more often when it hasn’t succeeded.
  • It must not attempt more shortcuts.

If any of those fail, it doesn’t ship. We’ll publish the results with the release.

Results at releaseStarting model Ysra SWE
Tasks solved on the first tryHigher is better
“Done” claims that weren’tLower is better
Shortcut attemptsMust not increase
Cost per solved taskLower is better
Bars are placeholders, not results. Real numbers publish with the release.

The Ysra model family

Engineering first. Then your language, your numbers.

Ysra SWE comes first. Every model after it has to earn its place with verified results.

In training

Ysra SWE

The software engineer.

Investigates real repositories, makes changes across files, runs the checks, and hands back work a person can review.

Planned

Ysra Fast

Quick by default.

A smaller model for everyday work that hands harder tasks up to Ysra SWE.

Planned

Ysra Arabic & Darija

Your language, fluently.

Arabic, Moroccan Darija, French and English — including how people mix them at work.

Planned

Ysra Finance

Every number checked.

Invoices, reconciliation, ledgers and VAT, where answers can be verified exactly.

Research

Ysra Reviewer

A second set of eyes.

A focused critic that spots gaps, risks and missing evidence before work comes back to you.

Plain facts

What we’ll say — and what we won’t.

No borrowed benchmarks and no “trained from scratch”. Just what it is, where it comes from, and how your data is treated.

What is it built on?
Ysra SWE is post-trained from a leading open-weight model. We didn’t pretrain a model from scratch — we’re teaching a strong one how Ysra works. We’ll share the full lineage at release.
Is my code used to train it?
Customer code is excluded from training unless its owner permits it. Anything we can’t confirm we have the right to use is left out by default.
What about secrets in the data?
Every record is scanned for keys, tokens, passwords and credentials. If we’re not certain a secret is gone, the whole record is rejected — not “probably cleaned”.
When can I use it?
When it beats its starting point on real work it has never seen. Join early access and we’ll tell you first.

Coming soon

Be first to work with Ysra SWE.

Early access goes to teams already handing real engineering work to Ysra. Tell us about yours.