A polished answer can be completely wrong. A successful test can support a very small conclusion. An attractive chart can make a weak assumption look settled.
My approach to applied AI research is to make those boundaries visible before the interface makes them easy to miss.
Two current projects put that discipline under pressure: Saturview, a football intelligence product, and Fleet Kernel, a prototype for more reliable agent execution. They solve different problems. They ask a similar question: what actually supports the result?
An assistant should not improvise the evidence
In Saturview, free-form model interpretation produced wrong-team answers during development. Making the prompt sound more authoritative was not an adequate repair.
The implementation moved toward a narrower contract: the provider selects from supported record identifiers, and the application renders the relevant trusted facts. Invalid or empty selections are rejected rather than dressed up as an answer.
That is not a claim that every answer is now complete. Relevant selection can still be imperfect. It is a decision to reduce the part of the job in which fluent language can hide a factual mistake.
The business version is straightforward. Before an assistant can explain your pipeline, inventory, or customer history, it needs a dependable relationship with those records.
A retrospective is not a prediction
Saturview keeps actual results, hypothetical seasons, and prospective picks distinct. A model selected after a historical period can be studied on that period, but the resulting score is not evidence that it predicted those games in advance.
That distinction matters even when a retrospective looks good. Changing the model after seeing the outcomes changes what the evaluation can support.
My preference is to preserve what was known, when it was known, and which configuration produced the output. A result without that context is too easy to overstate.
The current public beta is a functioning research product. It is not a promise of betting profit or a validated predictive edge.
Let code carry the contract
Fleet Kernel explores durable tasks, scoped execution, routing policy, budgets, and evidence-gated completion. The aim is to make important operating decisions inspectable rather than leave them entirely inside an agent conversation.
One implemented prototype behavior fingerprints a source before deciding whether work needs to run. When nothing relevant changed, it can stop before calling a model. That behavior is narrower and easier to test than a general promise that an agent will always spend efficiently.
Other parts of the prototype check whether submitted artifacts match the task’s acceptance conditions. A worker saying “done” is not the same thing as those checks passing.
Universal cross-harness production readiness is not established by the current prototype. The deployment boundary is part of the finding, not an inconvenient footnote.
Failed experiments still belong in the record
The current Astra Superharness project record includes an HTML benchmark that did not complete successfully through the plugin. That does not support a cost-saving claim. It identifies a limitation to investigate.
I would rather publish the boundary than turn an incomplete run into a victory lap. Research should improve the next decision, not protect the previous claim.
What this means for a business
Start with a task whose result you can inspect. Decide which facts must come from a trusted system, which decisions require approval, and what evidence counts as completion. Keep unsupported inputs visible. Test failures as deliberately as happy paths.
AI expertise is not just knowing which model to call. It is knowing how much confidence the surrounding system has earned.
Source note: Based on my project records for Saturview, Fleet Kernel, and Astra Superharness, reviewed September 20, 2026. These are implementation observations and research boundaries, not independent benchmark results.