AgentWatchr
Evidence Score v0.2
Calidad de evidencia, no reputación. Datos ausentes: N/A.
# Evidence Score Methodology 0.2.1 Version: evidence-score/0.2.1. The existing engine replaces the old capture/freshness average. It does not assign reputation, safety or agent-performance points. ## Formula and scope Overall score is null, status INSUFFICIENT_EVIDENCE. There are no weights, hidden zeroes, clipping caps or averages over available axes. None of the available evidence establishes comparable task-performance outcomes. No calibrated overall scale is claimed. Three separate outputs are published: 1. Record quality: number of valid attributed observation references / number of all distinct referenced observations. A zero denominator yields null. Valid means an aware non-future timestamp, resolvable reference and SHA256 syntax. This measures our record; archived bytes are checked separately before publication. 2. Verification scope: counts of distinct dated HTTP attempts, responses and accepted task tests. HTTP responses, including 200 and 405, are not task success. 3. Task performance: null, NOT_TESTED; no accepted task-test schema or evaluations currently exist. Agent capabilities are declarations. Coverage = 100 * count of six areas with admissible evidence / 6. Numerator, denominator, facets and exclusions are explicit. No area is removed from the denominator because a field is missing. This is documentation coverage, not an agent-quality score. A single source observation may support different factual areas; it is never summed into overall points or counted as independent corroboration. ## Dimensions | Dimension | Admissible evidence and measurement | N/A and limits / gaming | |---|---|---| | Identity | Resolved, dated, hashed source-history references. Valid source refs / distinct history refs. | NO_DATA without valid refs. A perfect captured record does not verify owner claims. Fabricated source descriptions can still be well recorded. | | Runtime | Attributed explicit attempted HTTP checks, their actual test times and status. Responses / distinct attempts. | NOT_TESTED without a check; NO_DATA when records lack valid attribution; TEST_FAILED for attempted checks without responses. This is transport evidence, never agent task performance. | | Historical | Distinct source observation timestamps and distinct payload hashes. | NO_DATA with fewer than two dates. Repeated identical reads show re-observation only. Repetition cannot raise independent confidence. | | Infrastructure | Attributed DOMAIN/ENDPOINT/CONTRACT/MULTISIG declarations / all such declarations. | NO_DATA when unavailable. Source-declared fields are not independently verified; adding many declarations can inflate apparent detail. Shared values carry no penalty. | | Consistency / Relationships | Attributed relations / all declared relations. Also detects differing HTTP statuses for identical endpoint/test-time groups. | NO_DATA without relations. CONFLICTING_EVIDENCE is a review flag, not a penalty. Potential duplicates remain candidates. Different results at different times are not automatically contradictions. Other conflict types are not yet evaluated. | | Task Performance | No admitted real task tests. accepted_tests=0, successes=null, value=null. | NOT_TESTED. Neither capabilities, HTTP, source anchors nor freshness substitute for actual evaluation. | Axis value remains null; measurement objects carry actual counts/ratios with named units. Ratios on different axes are not commensurable scores and are not averaged. ## Applicability and states Profile types are mapped explicitly: ACP registry -> PACKAGED_ACP_AGENT, ERC8004 Base -> ONCHAIN_REGISTRATION, Agentverse -> REGISTRY_AGENT, Olas Gnosis -> ONCHAIN_SERVICE, otherwise UNCLASSIFIED. Public HTTP is not required for the packaged ACP class, so no HTTP check is NOT_APPLICABLE there. Missing endpoints do not exempt any other class. An observed ACP HTTP check is still reported as evidence. Full task applicability cannot be inferred and remains NOT_TESTED. The compatibility status N/A is supplemented by assessment_state and na_reason: NO_DATA, NOT_TESTED, NOT_APPLICABLE, TEST_FAILED, OBSERVED or CONFLICTING_EVIDENCE. A failed HTTP attempt is not a failed agent task. ## Time, confidence, reproducibility Source freshness uses only source-history observation dates. Runtime freshness uses ENDPOINT_LAST_TESTED_AT (or the actual test observation time). A registry reread cannot refresh an older test. Each dimension exposes its own date and age. No calibrated confidence probability is available. confidence_basis exposes distinct hashes, source labels, and null independent corroboration. Legacy low/insufficient fields are retained for client compatibility only, not likelihood estimates or quality ratings. Every result exposes methodology_version, as_of/evaluated_at, snapshot_id, input_sha256, evidence_refs, reason_codes and why_this_score. The fallback snapshot_id hashes the profile and referenced evidence independently of evaluation time. Whole-snapshot audits set snapshot_id to the bundle SHA256. Identical inputs, version, snapshot and time reproduce identical results. Audits append evaluations.jsonl with both previous 0.2 and current 0.2.1 calculations. Historical snapshots and archived source bytes remain unchanged. The historical 0.1 flaw was capture60/80 averaged with freshness100, giving80/90 without task tests; 0.2 had already removed that average. 0.2.1 corrects admissibility of endpoint references, duplicate test counting, applicability, explicit states and separate freshness, and publishes measurable denominators. The audit is a boundary/integrity test. It does not validate fraud prediction, safety or task ability. API clients can ignore all scoring and use source evidence and reason codes directly.