
By Shaun Modi, CEO, Capitol AI
Test the workflow behind the demo
A generated diligence memo can make a platform demo look impressive with polished analysis, comprehensive citations, and a result delivered in a few short minutes. But will it hold up when your team has to vouch for it? The real evaluation needs to consider the time your team spends producing a finished deliverable that’s up to your organization's standards, including validating sources, correcting errors, and securing approval from stakeholders. That is the work a platform has to improve. Most evaluations fall short of that rigor.
McKinsey’s 2026 State of AI survey found that 80% of respondents saw gains in individual productivity, but only 37% reported a positive contribution to enterprise earnings before interest and taxes. The firms reporting the strongest results were also more likely to redesign their workflows. For enterprise, the practical question is whether a platform improves the whole process, including the work needed to check and approve the final deliverable.
Every organization is different, so off-the-shelf tools often need customization to fit an institution’s existing process. In our earlier article on building or buying enterprise AI, we explained that most institutions will do some of both. They will use models developed elsewhere while building around the knowledge and the expertise that make their own work valuable in the first place. The choice is which parts of that process you want the vendor to provide and which your team will build and maintain.
Three routes and six questions to help you choose the right partner
The three routes below are just starting points, and institutions may combine elements of each:
A specialized vertical tool, designed around a specific industry or even more narrow like a defined task. One example is financial research.
A general assistant from a lab such as Anthropic or OpenAI, working directly with the provider or with a consultant to adapt it to your institution.
An internal build, with your engineers developing and maintaining the system around hosted or open models.
Capitol fits between the general-assistant and internal-build routes: you buy the workflow platform and use it to build and standardize your institution’s processes across approved models. That still leaves a decision on how much your experts can configure, and how much will need vendor or engineering support. These six questions below test the division of responsibility.
Give each vendor the same representative case, including an exception, and compare time to an approved result, corrections required, and total cost. Then ask your team to change one step and run it again.
1. Can we customize the workflow ourselves?
Your workflow needs to follow your institution’s standards, from inputs and calculations to presentation. Ask who can change the analytical method, permitted sources, calculations, and review steps. Does each change require a support ticket, a consultant, or your own engineers? In a diligence workflow, try changing how a specific number is calculated, and make sure the vendor allows you to drill in that deeply. Watch who makes the change, how it is tested, and whether it affects all the connected steps. You will be able to spot points of friction very quickly.
2. Where does our data go, and who controls it?
Establish where data is stored and processed: the vendor’s cloud, your private cloud, or your own infrastructure inside the established perimeter. Then follow an actual model call. What information leaves your environment, what is filtered or redacted, and what is retained or used for training? Hosting the application internally does not establish where every model call goes. Ask to see the configuration and terms that meet your requirements. If you need someone technical with you, bring them in.
3. Who can build, change, approve, and run workflows?
Look for distinct roles for administrators, workflow builders, and users. An analyst running an approved process should have only the authority their role requires. Ask how changes are approved, where human review occurs, and whether the workflow waits for that approval before continuing. User access alone does not define these controls.
4. Does our institutional knowledge survive the person who created it?
In a general assistant, some of your analysts’ best work stays siloed in their chat history. Your approved analyst’s processes should become reusable instructions, source rules, evaluation criteria, and review steps for everyone else who needs to produce a similar output. Ask another employee to run the workflow without its author coaching them. Can they inspect the method and its version history? Check what workflow definitions and records you can export if you later change platforms.
5. Can we verify the result and catch mistakes?
Mistakes erode trust with clients and regulators. Pick a material claim and trace it to the source and calculation behind it. Ask how the output is tested against your standards and what happens when it fails. A citation does not guarantee a correct interpretation, and a repeatable workflow does not guarantee a language model will return the same result to the same prompt each time. The system should make exceptions visible and route them for review.
6. Can we swap models and understand the cost breakdown?
Ask whether different steps can use different approved models, or ordinary code where a model is unnecessary. A simple lookup does not need a frontier reasoning model. Ask what must be retested when a model changes. When thinking about costs, think in terms of outcomes, not tokens. Compare the cost of a completed, reviewed deliverable, including setup, ongoing support, model usage, retries, and human review. A subscription price alone will not tell you what the workflow costs to operate.
Put the framework to work
In Capitol’s Composer, connected steps can combine data retrieval, calculations, model calls, evaluations, and human review. Your experts define the methods and quality standards, with controls built into the workflow.
Bring a workflow to a Capitol demo. We will show how your data, process, and review requirements fit together so you can evaluate the system against the work your team actually needs to do.
Related reading: What Is Sovereign Agentic AI? and Renting Intelligence vs. Building Intelligent Structure.
Frequently asked questions
What does claim level traceability need to show?
It should connect an individual assertion to the evidence and workflow steps supporting it. For a calculated result, that includes the relevant inputs and calculation. A traceable claim still needs review: a real source can be misread, and a correctly cited figure can be used in an incorrect analysis.
Does model independence make workflows portable between vendors?
Not necessarily. The ability to change a model within a platform is different from the ability to move the workflow to another platform. Check export formats, runtime dependencies, record retention, and contractual rights separately.
How can we test repeatability without expecting identical AI output?
Run representative cases against defined standards for accuracy, completeness, source use, and required approvals. Compare results across users and after controlled changes. The aim is consistent adherence to an approved process and quality threshold, with exceptions visible for review.
What should you compare when evaluating enterprise AI platforms?
Compare six things on one real piece of work: control over the method, who can build and approve a workflow, the evidence behind a claim, the data path into each model call, cost per run, and what you can take with you. Establish who can change the analytical method and how long that takes, whether a material claim traces back to its source and the step that produced it, what information reaches an external model, what a single run costs broken down by step, and what exports exist if the relationship ends. A prepared vendor sample will not test any of these.


