The short version
Key takeaways
- Begin with critical scenarios and failure costs.
- Separate verified capability from roadmap statements and assumptions.
- Include adoption, administration, integration, security, and exit costs in the score.
Define outcomes and constraints
Write the operational result first: reduce duplicate order entry, shorten proposal-to-start time, create a reliable customer handoff, or give managers one view of aging work. Add constraints such as budget, implementation window, data residency, accessibility, user count, existing systems, and regulated records.
Choose three to seven critical scenarios. For each, describe the starting event, people involved, information needed, successful end state, exceptions, and evidence created. These scenarios become the demo script and acceptance test.
Weight criteria by business impact
| Dimension | Example weight | What to test |
|---|---|---|
| Workflow fit | 30% | Critical scenarios and exceptions |
| Implementation | 15% | Configuration, migration, training, ownership |
| Integration and data | 15% | APIs, exports, identity, source of truth |
| Security and control | 15% | Roles, logs, retention, recovery, vendor practices |
| Usability and accessibility | 10% | Representative users and devices |
| Commercial fit | 10% | Total cost and contract terms |
| Exit readiness | 5% | Data export, deletion, transition assistance |
Weights should reflect consequence. A system that stores signed agreements should weight evidence and retention differently from a lightweight brainstorming tool.
Use an evidence scale
Score each requirement from zero to four: unavailable, claimed, demonstrated, tested with your scenario, or proven in a controlled pilot. Record the evidence link, date, product edition, limitation, and evaluator. A feature shown in a slide should not receive the same score as a workflow your team completed.
Label roadmap items separately. Do not purchase a current product primarily for an uncommitted future capability. When a vendor says “supported through integration,” identify the specific connector, owner, additional subscription, sync direction, failure behavior, and support boundary.
Calculate total operating effort
Include licenses, usage, implementation, migration, customization, training, administration, integration maintenance, support, change management, and expected rework. Estimate internal hours as well as cash. A lower subscription can be expensive if it creates manual reconciliation every week.
Document which team owns configuration after launch. If every small change requires a consultant, that is a commercial dependency. If every user can change critical rules, that is a control risk.
Run a reversible pilot and decide explicitly
Use representative data with appropriate privacy controls. Test normal work, peak volume, permissions, corrections, exports, support, and an outage or rollback scenario. Interview users about where they hesitated or created workarounds.
Summarize the recommendation, evidence gaps, risks accepted, mitigation owner, success measures, and a review date. A scorecard supports judgment; it does not replace it. A product can score highest and still be wrong if a single non-negotiable requirement fails.
Common questions
Frequently asked questions
How many vendors should reach the pilot stage?
Usually the smallest set that can plausibly meet the non-negotiable requirements—often two or three. Broad demos consume time without improving evidence.
Should price be scored separately?
Yes, but use total cost over a realistic period and distinguish predictable subscription cost from usage, services, migration, and internal operating effort.