← Evals Console

Experiments

Each compares a baseline with our method on the same cases. Improvements are claimed only when 95% intervals don't overlap; a loss is reported as a loss.

Grounding

measured

Hypothesis: Claim-level evidence with per-claim verification produces fewer unsupported factual claims than document-level citations.

Baseline
generate@baseline: same facts, sources listed once per email
Method
generate@v1: every sentence cites its fact ids; verify.ts with one retry
Data
3 accounts × 10 runs × 2 arms on the seed fact history
Pass bar
Method's 95% interval for unsupported_claim_rate sits entirely below the baseline's.
Pre-registered
experiments/grounding.md

Labels in experiments/grounding.labels.json were drafted by the coding agent and haven't been reviewed by a human yet. Treat results as provisional.

Not run yet.

Unsupported claim rate

Not run yet

    Judge agreement (verify.ts)

    Not run yet

      Runs from the terminal (paid model calls): LLM_MODE=live pnpm exp grounding

      Feedback learning

      measured

      Hypothesis: Routing each edit or rejection through the taxonomy sends more feedback to the fix that addresses it than sending everything to prompt review.

      Baseline
      approve/reject only; every negative signal goes to prompt review
      Method
      fact diff + edit classifier + routeFeedback
      Data
      17 labeled cases (10 edits, 7 rejects)
      Pass bar
      routing_accuracy ≥ 13/17 with intervals separate from the baseline's.
      Pre-registered
      experiments/feedback.md

      Labels in experiments/feedback.labels.json were drafted by the coding agent and haven't been reviewed by a human yet. Treat results as provisional.

      Not run yet.

      Routing accuracy

      Not run yet

        Graph corrections

        Not run yet

          Classifier accuracy

          Not run yet

            Autonomy

            simulated

            Hypothesis: At zero critical auto-executions, risk.ts needs review on fewer actions than approving everything; confidence-only gating can't reach zero criticals.

            Baseline
            approve everything; confidence-only (> 0.8)
            Method
            risk.ts policy at trust level 3
            Data
            60 labeled synthetic actions
            Pass bar
            critical_auto_executed = 0 and lower human_burden than approve-everything.
            Pre-registered
            experiments/autonomy.md

            Labels in experiments/autonomy.labels.json were drafted by the coding agent and haven't been reviewed by a human yet. Treat results as provisional.

            Not run yet.

            Human burden

            Not run yet

              Critical actions auto-run

              Not run yet

                Reference retrieval

                designed

                Hypothesis: Ranking past drafts by similarity × validated outcome × evidence quality lowers the edit rate compared with similarity alone.

                Baseline
                similarity only
                Method
                similarity × validated outcome × evidence quality
                Data
                needs live reps
                Pass bar
                Edit rate per draft lower, with separated 95% intervals.
                Min. sample
                200 decided drafts per arm
                Pre-registered
                experiments/references.md
                Needs live reps producing real decisions.

                Attribution

                designed

                Hypothesis: Accounts that get workflow follow-ups reply and book meetings more within 14 days than matched holdout accounts.

                Baseline
                last-touch attribution
                Method
                holdout accounts / matched comparison
                Data
                needs real outcomes
                Pass bar
                14-day reply + meeting rate higher, with separated 95% intervals.
                Min. sample
                300 accounts per group
                Pre-registered
                experiments/attribution.md
                Needs real outcomes over weeks.