Istanbul-based Loxi scored 58.2% on JobBench, leaving GPT-5.6 SOL behind and drawing level with Claude Fable 5. Its model is not a frontier system, its product is free, and it runs at under a tenth of the cost. The result points to a shift the AI industry has been slow to acknowledge: harness engineering and domain-specific training now matter as much as raw model scale.
—
Washington University researchers built JobBench around a different question than most AI benchmarks ask. Instead of measuring how much economic value could be automated, it starts from 1,500 professionals across 35 occupations rating which of their own tasks they would actually hand to an agent. The 130 tasks that came back are the work professionals want to delegate. Each deliverable is scored against an average of 36 binary, expert-written criteria, and no credit is given for a right number reached through wrong reasoning.
On the main split of that benchmark, Loxi scored 58.2%. That puts it level with Claude Fable 5 (57.4) and ahead of Kimi K3, Qwen 3.8 Max, Claude Opus 4.8 and GPT-5.6 SOL, which finished at 45.4. Every one of the 65 tasks produced its deliverables. No time-outs, no refusals, no empty folders.
The model under Loxi is not a frontier model. It is a small model the team trained specifically for knowledge work. Loxi is free, and the beta always runs the latest experimental build. What the company sells, in effect, is the layer between the model and the user.
That layer is where the score comes from.
Where the score comes from
Between the model and the user sits the harness. It decides how a task is planned, how files are read and produced, how the agent checks its own work, which skills it loads for a spreadsheet versus a legal memo, and the isolated environment where its code runs. For the past eight months, that layer has been the whole job at Loxi: thousands of controlled experiments on model and harness fit, one change at a time, measured on real tasks.
Independent research supports the claim that this layer carries more weight than the leaderboard suggests. LangChain changed only its agent runtime and moved from thirtieth place to fifth on Terminal Bench 2.0, raising its score from 52.8% to 66.5% without touching the underlying model. A paper published on arXiv found that holding the base model fixed and redesigning the harness lifted pass@1 from 69.7% to 77.0%, beating a human-designed Codex harness in the process.
As the Loxi team puts it, the same class of model, the same tasks and the same judge produced a different score once the harness moved. Benchmark tables tend to credit the model for work the harness is doing.
Domain-specific models and the changing cost math
Gartner expects more than half of enterprise generative AI deployments to be domain-specific by 2027, up from roughly 1% in 2024. The firm also projects that organizations will use small, task-specific models three times as often as general-purpose large language models by that year.
The evidence is already visible in finance and law. Bridgewater Associates worked with Thinking Machines Lab to fine-tune a model on how its own investment managers actually work, and the result outperformed frontier models at one-fourteenth of the cost. Harvey AI fine-tuned the Kimi 2.6 model on a legal benchmark and gained roughly 40%, reaching frontier-level performance at one-eleventh of the cost.
Loxi follows the same logic. JobBench tasks are public, the product is free, and user feedback flows directly into model and harness optimization. The capability is easy for anyone to check.
Turkey’s AI push meets a product moment
Loxi is built by an Istanbul-based team, and its timing lines up with a broader institutional push in Turkey. The country’s National AI Action Plan for 2026 to 2030, issued by presidential decree in August 2026, frames AI as a matter of digital sovereignty and organizes policy around four pillars: perceive, benefit, produce and govern. It sets targets of 1 GW of data center capacity, at least $10 billion in private investment, AI literacy training for 5 million people and 10,000 advanced specialists by 2030.
The plan also calls for developing and exporting sector-specific language models in priority areas such as health, energy and smart manufacturing. Industry Minister Mehmet Fatih Kacır has said Turkey intends to be among the leaders of the global technology transition rather than a spectator to it.
The startup ecosystem has been moving in the same direction. Turkey reached eight unicorns in 2026, and the number of candidates in the Turcorn 100 program nearly doubled from 23 in March 2025 to 43 in the first half of 2026. Ankara-based Herdr raised $6 million in seed funding to build an open-source AI agent runtime for terminals. The country’s stated goal is 100 unicorns by 2030.
Read the latest results on the [Loxi news page](https://loxi.works/news).
Beyond the two-bloc map
A July 2026 Bank of America report describes an AI race concentrating around the United States and China. The US holds its lead in private investment, advanced chips and compute infrastructure, while China leverages manufacturing scale and low energy costs. Outside those two, BofA names South Korea as the strongest candidate, followed by the UAE, with Canada, Germany, Israel, the Netherlands, Singapore, Switzerland and the UK forming a second tier.
That map leaves little room for anyone else, and it is the reason Loxi’s result is worth noting. A country’s position in AI is not determined only by infrastructure and capital. Software engineering depth and systems design can move a small team into the same performance band as the largest labs, and a two-bloc framing does not capture that.
What it means for work
PwC‘s Global AI Jobs Barometer, published in June 2026, found the labor market splitting into two tracks. Specialized roles, where AI automates routine work and pushes human judgment to the fore, are growing twice as fast as democratized roles and carrying 42% faster wage growth. Companies that use AI most effectively report 52% employment growth against 36% for the least exposed, and 24% wage growth against 17%.
JobBench measures the work at the center of that shift. The tasks are routine but hard, repetitive but demanding of close attention, and they are the tasks professionals themselves said they would hand off. An agent that can do them reliably is not a replacement for a person. It is a tool that lets a person do more of the part that requires judgment.
The next phase
Loxi’s results show that model size is not the only variable that decides who competes at the top. A small team, a model trained for knowledge work and eight months of disciplined harness engineering reached the same performance band as the largest labs, at a fraction of the cost.
That is not only a benchmark result. Domain-specific models are rising, harness engineering is becoming a first-class discipline, and players outside the US and China are starting to appear on the board. Turkey’s first AI products aimed at global benchmarks are one example. The next phase of competition will be decided less by who trains the largest model and more by who builds the best system around it.



