iTnews Asia
  • Home
  • Features
  • Partner Content
Partner Content

TRACES puts AI to the test for the next frontier of AI Discovery

TRACES puts AI to the test for the next frontier of AI Discovery

By iTnews Asia Team on Sep 4, 2026 10:05AM

AI is permeating every industry but for the scientific R&D community, mainstream generative AI tools won’t cut it. What the community needs is not a technology that packages and touches up answers – it needs something that helps researchers test hypotheses, make use of evidence, and chart the most plausible way forward when the final answer is not known.

Apodex was founded to provide this capability that generalist AI tools today do not have.

iTNews Asia speaks to Brian Wang, AI Research Scientist, Apodex, to find out more about the company and its TRACES AI benchmark tool.

iTNews Asia: For those not familiar with Apodex yet, could you introduce the company and explain what TRACES stands for?

Wang: Apodex was built on a founding thesis hypothesising that scale and data alone will not get AI from pattern matching to real discovery. Getting there takes new ideas from neuroscience, information theory, physics and formal verification, along with people who know how to bring those ideas together.

That belief is what led us to build what we call Discoverative AI, systems designed to find things nobody knows yet. TRACES is the benchmark we developed to evaluate this kind of AI. The name reflects the six capabilities we test for: Tools, Repair, Alternatives, Coherence, Evidence and Scope. Together, they capture what it actually takes to investigate real, verifiable solutions.

iTNews Asia: What is TRACES and what problem is it designed to solve for AI development?

Wang: TRACES is a reality benchmark testbed for a different, more complex class of AI work that is rooted in the unknown. These are problems where a system or algorithmcannot simply look up the answer because nobody has established it yet.

A typical benchmark asks an AI to answer a question, write some code or solve a defined task. TRACES asks the system to operate more like a junior researcher or technical investigator. It receives a relevant working environment and has to decide how to proceed within it, deciding on which tools to use, what evidence to collect, which hypothesis to test and when to change course.

The final output depends heavily on this methodology. Did the system run the right analysis? Did it use the output correctly? Did it test alternatives? Can an expert trace a claim back to the underlying evidence? These are the behaviours that genuine discovery work depends on.

iTNews Asia: Why does Apodex believe the industry needs a new benchmark for scientific discovery?

Wang: Because the work itself changed, and the way we measure it did not. Until recently, AI was mostly asked to answer - summarise this, write that, solve a problem that already has a solution. Increasingly it is asked to investigate: to work for hours or days inside a real environment, choose its own experiments, and come back with a conclusion nobody has verified yet. That is a different job, and the tests the industry inherited were built for the previous one.

There is a structural reason conventional benchmarks cannot follow it. A benchmark that can mark your answer must already hold that answer — which means, by construction, it is testing a problem someone has already solved. That is useful for measuring how quickly a model reaches a known result. It cannot tell you whether a system would find something that is not yet known.

TRACES was created to make discovery processes and their failure modes verifiable and traceable. We want developers to be able to see whether their systems are progressing from fluent generation toward disciplined investigation, and to understand exactly where that progression breaks down.

iTNews Asia: Who is TRACES designed for, and how can users and industry enterprises use it effectively?

Wang: There are two main groups. First and foremost are the AI teams building agents, tool-use systems, solver frameworks or advanced models. These workloads require pinpoint-accurate diagnostic evidence about whether their systems can sustain a complex investigation over many steps.

The second is domain experts and organisations with difficult problems of their own. This could be a research group, a life-sciences company or an industrial enterprise. They may have data, tools and a high-value question, but need a rigorous way to evaluate whether an AI system is genuinely helping solve it.

For both groups, the value of TRACES lies in its creation of a pioneering standard for assessing capability in settings to discover the unknown.

iTNews Asia: Where do the real-world problems behind TRACES come from?

Wang: TRACES was built to test AI against problems with the potential to unlock real breakthroughs, rather than against questions designed simply to produce a benchmark score.

We began with a two-month scan across 561 industries in 16 sectors, building a registry of 423 high-value problems where progress is held back by the same constraints researchers face every day. In some cases, evidence may be incompleteor feedback loops slow or expensive. Several explanations may appear credible at once as the work may depend on specialist data or scientific tools.

From that registry, we selected and developed the first 20 TRACES environments across frontier-model development, biomedical discovery, clinical translation, scientific engineering and deployment-oriented intelligence. Each one turns a meaningful problem into a working setting where an AI has to do more than give an answer.

iTNews Asia: When does TRACES become most valuable for organizations assessing AI systems?

Wang: TRACES is most valuable when the cost of trusting a plausible but wrong AI answer is high. The moment AI begins moving from explaining what we know to helping uncover what we do not, TRACES becomes essential. That is where conventional benchmarks fall short and the final answer alone cannot tell you whether an AI used sound methods, learned from evidence or simply arrived at something that looks convincing.

That is often the case when feedback is delayed or incomplete, for example, when a proposed candidate may only be tested later in the physical world or when a technical report needs to be audited before it can inform a decision. In these settings, waiting for an eventual outcome could be a business-critical risk.

For AI developers, this is also the point at which scaling alone may stop being an adequate measure of progress. As systems are expected to take more initiative, their ability to repair, reassess and remain evidence-led becomes as important as their ability to generate.

iTNews Asia: How does TRACES evaluate an AI system differently from conventional benchmarks?

Wang: Most existing AI benchmarks work like a standardised exam. They give a model a question, then mark the final answer. TRACES works more like a researcher observing another researcher at work. We give an AI a reality-based environment, which may include relevant data, literature, code or specialist tools. The model itselfmust decide what to investigate with the tools available and adjust its approach when new evidence points in another direction.

That is why we value assessing and verifying both the outcome and the working record an AI leaves behind. In total, we evaluate six behaviours that are essential to sustained investigation: tool use, error repair, consideration of alternatives, coherence across a long chain of work, evidence discipline and clear limits on what a conclusion does and does not establish. This translates once previously abstract principles into task-specific checks and repair loops for each environment.

Each TRACES assessment is anchored to recorded actions and clear scoring criteria, with independent review where needed. This gives developers and researchers the trust and confidence to determine whether a system's conclusion was genuinely earned and where it needs to improve.

To reach the editorial team on your feedback, story ideas and pitches, contact them here.
© iTnews Asia
Tags:
apodex partner content

Related Articles

  • Why network’s next bottleneck is not bandwidth, but the route our data takes
  • Agentic AI tackles RTL verification’s productivity gap
  • How tech leaders can govern AI outcomes at scale
  • AI voice agents and the human touch: A new playbook for SME customer engagement
Share on Twitter Share on Facebook Share on LinkedIn Share on Whatsapp Email A Friend

Most Read Articles

TRACES puts AI to the test for the next frontier of AI Discovery

TRACES puts AI to the test for the next frontier of AI Discovery

Why network’s next bottleneck is not bandwidth, but the route our data takes

Why network’s next bottleneck is not bandwidth, but the route our data takes

Agentic AI tackles RTL verification’s productivity gap

Agentic AI tackles RTL verification’s productivity gap

TNB's One-Stop-Centre meets hyperscale data centres' power demands in Malaysia

TNB's One-Stop-Centre meets hyperscale data centres' power demands in Malaysia

All rights reserved. This material may not be published, broadcast, rewritten or redistributed in any form without prior authorisation.
Your use of this website constitutes acceptance of Lighthouse Independent Media's Privacy Policy and Terms & Conditions.