Evidence & Validation
What we measure, what would disprove us, and what happens if we're wrong.
The Commitment
CTE publishes pre-defined validation gates at Day 30, 60, and 90 to test whether the Enterprise Agent Architecture thesis — that enterprises need a vendor-neutral, execution-time governance layer for their agent workforce — is supported by real, external signal. If a gate fails, we say so publicly and revisit the thesis, not the scorecard. The gates measure demand, authority, and pipeline.
If this succeeds, we'll scale it.
If it fails, we'll say so.
These are the gates that decide.
Validation Gates
The gates are pre-defined falsifiability criteria: Day 30 tests external demand for the open-core work. Day 60 tests whether the work earns authority — citation and engagement from the CIO/CISO/architect audience. Day 90 tests whether interest converts to an active design-partner pipeline. Failure on any criterion triggers a public reckoning, not a quiet reframe.
We've defined specific, measurable criteria at three checkpoints. Failure on any criterion triggers the halt path. No exceptions, no reinterpretation.
Is anyone actually pulling the open-core work — is there external demand without us pushing it?
- Open-Core Demand constitutional-agent sustains > 200 installs / month on PyPI, unprompted
- Registry Pull the published dataset and package are pulled organically, with no promotion driving it
- Qualified Inquiry at least 1 qualified design-partner inquiry from the free offer
If failed: no external demand — the premise that enterprises want vendor-neutral agent governance is unsupported at this stage. Publish the null result; revisit the wedge.
Is the work earning recognition in the category — cited, referenced, engaged by the people who would actually adopt it?
- Citation at least 1 external citation or reference to a published preprint (DOI)
- Practitioner Engagement substantive engagement from the CIO / CISO / enterprise-architect audience — replies and threads, not vanity impressions
- Category Signal the framing shows up in others’ work or analyst commentary
If failed: reach without recognition — the category position is not forming. Publish what landed and what did not; sharpen the position or concede it.
Does interest convert into an actual design-partner conversation — the one success signal that is not vanity?
- Active Conversation at least 1 active design-partner conversation underway
- Reach → Pipeline a measurable, attributed path from content or citation to inquiry
- Runway survival intact (runway ≥ floor) so the thesis gets a fair test
If failed: distribution works but the offer does not — fix the offer, not the content. Publish the conversion gap; decide persist-or-pivot on the evidence.
What Success Looks Like
If we pass all gates:
- The open-core work has real, unprompted external demand
- The Enterprise Agent Architecture framing is cited and engaged by the CIO/CISO/architect audience
- At least one active design-partner conversation is underway
- We have evidence to justify deepening the framework and its reference implementation
- The published preprints and open-core artifacts are contributing to the agent-governance field
What Failure Looks Like
If we fail any gate:
- We publish exactly what we learned (including why it failed)
- We halt the current distribution thesis publicly and transparently
- We do NOT quietly reframe the failure or move the goalposts
- The analysis and artifacts remain public for others to build on
Why We're Publishing This
Most AI ventures launch with bold claims and vague success metrics. If they don't work, they quietly pivot or shut down. Nobody learns anything.
We think that's backwards.
By publishing our validation gates in advance, we're committing to a specific, falsifiable hypothesis about enterprise agent governance. If we're wrong, the world learns something. If we're right, the evidence is credible because it was defined before we knew the outcome.
This is what research-first actually means.
Published Framework
Our frameworks are publicly documented and citable:
Saleme, M. K. (2026). Enterprise Agent Architecture: The Case for a Fifth Architecture Domain for the Agentic Enterprise. Zenodo. https://doi.org/10.5281/zenodo.21105314
Saleme, M. K. (2026). Authorized but Composed: Cross-Session Risk Composition as an Agent-Governance Control. Zenodo. https://doi.org/10.5281/zenodo.21400261
Cognitive Thought Engine. (2026). Constitutional Self-Governance for Autonomous AI Systems. Zenodo. https://doi.org/10.5281/zenodo.19162104
Saleme, M.K. (2026). Detecting Normalization of Deviance in Multi-Agent Systems: Empirical Evidence for Graph-Based Behavioral Drift Detection. Zenodo. https://doi.org/10.5281/zenodo.19195516
Saleme, M.K. (2026). Beyond Identity Governance: A Protocol-Level Security Testing Framework for Multi-Agent AI Systems. Zenodo. https://doi.org/10.5281/zenodo.19343034
Saleme, M.K. (2026). Community-Driven Security for AI Agents: Evolution of an Adversarial Testing Framework. Zenodo. https://doi.org/10.5281/zenodo.19343108
These preprints establish our theoretical foundations, component definitions, and validation approach. We publish methodology before validation so our framework can be scrutinized independently.
See how your agent deployments hold up under governance pressure — a short, free self-assessment.
Take the Assessment Learn the Method