Before you start
You need:- An organization workflow you can edit, with steps that gather evidence and resolve the company or deal being assessed.
- A scoring framework in your organization, or permission to create one.
- An API key or a connected Metal MCP client.
read:workflows. Creating a family or configuring workflow scoring requires both read:workflows and write:workflows, plus the relevant app permissions. Reading scores requires read:scores and the target resource’s read scope, such as read:companies or read:deals.
How scoring fits together
The app calls a scoring definition a scoring framework. The versioned API object used by workflows is a Scoring Family, distinct from the legacyScoringFramework object.
A workflow pins a Family and rules version. Its scoring step evaluates measures, and publication mappings connect those outputs to a stored assessment. Published assessments attach to a company or deal, with an optional screening reference.
Selecting a framework does not score anything by itself. You still need evidence, scoring instructions, structured outputs, and publication mappings. A number in a report or PDF is not automatically a stored assessment.
1. Define the decision and rubric
Write the decision first. For example, “Should this deal advance to analyst review?” tells you which measures and evidence you need. For each measure, define:- What it assesses and which evidence supports it.
- What the score levels mean, including intermediate values.
- Whether higher means better or worse. Keep direction consistent within a blend; normalization does not reverse a risk scale.
- The numeric range and increment. Use
step: 1for whole numbers, or a smaller step when the rubric permits decimals. - How to handle missing, stale, or contradictory evidence, and when a person must review the result.
instructions field.
Use requiredValues for results every published assessment must contain. A required blend needs all of its component measures. If you set a primaryValue, it must also be required.
All blend weights must be positive and total one. min and max ignore weights when calculating the result, but still require them, and all component measures must share the same minimum and maximum. Blends reference measures, not other blends. Keep unsupported formulas, caps, and vetoes explicit in your decision logic rather than approximating them with a weighted average.
Create or reuse a Family
Uselist_scoring_families and get_scoring_family to inspect an existing definition. To draft a new one, pass the following payload to propose_scoring_family. It validates the draft without saving it. Once reviewed, pass the validated payload to create_scoring_family, or use POST /v1/scoring-families with the same body.
append_scoring_family_rules creates a new immutable version; workflows pinned to an older version do not switch automatically.
2. Gather evidence and resolve the target
Collect the evidence and resolve company or deal identity before scoring. A project name, company display name, CRM ID, or document title is not a Metal company ID. Resolve the existing record rather than asking the scoring step to infer identity from a name. Give the scoring step the evidence, investment criteria, source dates, and references. Do not ask it to silently fill gaps from memory. Publication requires at least one input reference object with stringid and type fields, such as a screening or document reference. A document reference uses type: "data_entity". Preserve these objects in step outputs. Input references identify what was assessed; citations identify the passages that support a claim.
If the company or deal is unresolved, do not create a record merely to save a score. A configured target mapping can resolve to a missing value or null, leaving the assessment pending. The mapping must still reference a valid step, and configuration must include a company or deal mapping. A screening alone is not the target.
Pending assessments do not appear in current scores, score history, or benchmarks. Use list_pending_workflow_scores to find them, then publish_pending_workflow_score to bind the correct company or deal, or dismiss_pending_workflow_score with a reason. These are one-way transitions. Publishing or dismissing requires read:workflows and write:workflow_runs; binding also requires the selected target’s read scope and access to the source run.
3. Pin the rules and instruct the scoring step
SetscoringConfig to the Family ID and a positive rulesVersion. You can configure multiple Families, but each Family can appear only once. Families are organization-scoped and cannot be attached to global workflows.
Each run snapshots its scoring definitions. Include them in the scoring step’s prompt with:
outputSchema. Values and confidence are numbers, reasoning is a string, and citations are arrays of citation objects or handles. Include a boolean readyToPublish and a string array missingEvidence. Allow unavailable measures to be omitted so the schema does not force the agent to invent a number.
A completed assessment’s structured output could look like this:
4. Configure publication mappings
UsePUT /v1/workflows/{id} with the body below. For the MCP configure_workflow_scoring tool, add workflowId alongside scoringConfig.
Replace the illustrative Family ID with your returned ID and use its actual version. The referenced steps must already exist with matching output schemas:
resolvereturns acompanyIdstring.collectreturns aninputRefobject with required stringidandtypefields.scorereturns the structured assessment shown above.
output.results.structured for non-tool steps, or output.results for tool steps. Do not prefix them with results.structured. The publication condition is a CEL expression and uses the full steps[...].results... path instead. Mapping sources can be agent, completion, tool, or executeCode steps.
Map measures only. overall is absent from publication.values because Metal derives blends. Publication validates required values, numeric ranges, step alignment, confidence, and required reasoning. It also produces normalized values from 0 to 1; normalization is not a probability estimate.
Optional mappings include target.dealId, target.screeningId, event-level citations, and string-valued context. Each inputs entry maps one reference object, not an array. To score iterator items, set iteratorStepId and map that iterator’s child step outputs.
5. Use scores in a decision
Stored assessments are published during workflow finalization, not immediately after the scoring step. A branch or human-review step in the same run must use that run’s step outputs. Do not fetch the current stored score and assume it belongs to the run in progress. If routing needs a blend before publication, calculate the pinned formula in a deterministic code step. For this example,70 * 0.7 + 60 * 0.3 = 67. Do not ask the agent to invent an aggregate.
Use a branch condition for the business decision and publication.condition to decide whether to save the assessment. A rejected deal may still deserve a stored score. Conditioning publication on an “advance” verdict would lose that history.
An omitted publication condition means publish. A false condition skips that Family’s publication, which prevents saving incomplete assessments in this example. A condition evaluation error fails finalization rather than silently publishing.
Agree on thresholds and review requirements with your organization. A high score does not authorize CRM writes or bypass a human approval step.
6. Validate and read the results
Review examples of a clear fit, a clear non-fit, a borderline case, missing or conflicting evidence, and an ambiguous target. Check a retry for duplicate assessments and a rules update for unchanged older run definitions. Start with a workflow dry run to validate the configuration without saving Scoring Events. Inspect your workflow’s tools and code before assuming all side effects are suppressed. Then verify a controlled live assessment with the intended target and sources. A Family created through MCP starts inactive. Saving an event and making it the current result are separate actions. Confirm the Family is active before expecting its results in current-score reads, and coordinate activation with someone who can manage scoring frameworks in your organization. If replacing a legacy framework, agree on activation and rollback separately from authoring the new rubric. Useget_score_history to inspect stored assessments and get_current_scores to verify the selected result. Both require exactly one of companyId or dealId. For REST, use GET /v1/scoring/history?companyId={companyId} and GET /v1/scoring/current?companyId={companyId}, or the equivalent dealId filter. Follow metadata.nextToken when paging through history.
Next steps
Run and monitor workflows
Start runs, validate with dry runs, and inspect the results.
Scoring MCP tools
Read tool scopes and available Family and score operations.

