AI deployment, evaluation & accountability systems

Applied AI safety and accountability systems

Healthcare deployment experience, authorization controls, evaluation protocols, and production architecture for high-stakes AI — from EHR decision-support to deterministic evidence verification.

Deployment safety and authorization systems

Authorization states, reconsideration loops, and audit ledgers for institutional AI deployment — ensuring decisions remain explainable, challengeable, and correctable in production.

Judgment Commitment →

DeploymentGovernance

A proposed protocol for recording when an AI system's conclusions are reliable enough to act on. Its specification covers proposition registration, evidence-state freezing, and revision obligations.

Problem
AI systems make probabilistic claims. Institutions need deterministic commitments they can rely on and later revise.
Mechanism
Frozen evidence-state registration with explicit scope, authorization conditions, and revision obligations.
System requirement
Binds model inference to explicit evidence snapshots and mandatory revision triggers.

Authorization states →

GovernancePolicy

A graduated permission model for AI deployment — from prohibited through constrained use to routine reliance. Each transition requires explicit evidence and institutional sign-off.

Problem
Binary approval/rejection fails for AI systems whose risks emerge over time.
Mechanism
Four states: Not authorized, Observed, Constrained use, Routine reliance — each with specific conditions.
Design intent
A staged authorization model for increasing reliance only as evidence, monitoring, and correction capacity mature.

Reconsideration loop →

OperationsSafety

Six-step protocol for responding when new evidence, performance degradation, or escalation triggers call a prior authorization into question.

Problem
AI systems operate continuously but the evidence they rely on does not stay current.
Mechanism
Detect → Connect → Triage → Review → Decide → Propagate. Each step has defined inputs, actors, and outputs.
Operating control
Prevents silent failures when models update, context windows drift, or clinical guidelines shift.

Escalation design →

SafetyHuman oversight

Five conditions that force human review: patient-safety signal, performance degradation, new evidence, repeated overrides, and model/prompt/data changes.

Problem
Automatic escalation is either too sensitive (desensitizes reviewers) or too permissive (misses failures).
Mechanism
Trigger conditions paired with response levels — advisory notice, mandatory review, or automatic override.
Safety mechanism
Enforces mandatory human review or automated overrides before high-stakes execution.

Technical systems

Open-source systems and prototypes for evidence verification, capability tracking, and forecasting.

NextConsensus →

ForecastingEvidenceArchitecture

Forecasting platform and accountability protocol under development. It is designed to register frozen-evidence probability forecasts for observable medical-guideline, regulatory, and coverage transitions.

Protocol design
Point-in-time evidence state, immutable proposition registration, adjudication rules, probability history with Brier-score resolution.
Non-AI component
Deterministic observation layer (Refract) ingests structured change events from public sources. No model involved — pure verification and provenance.
Architectural separation
Isolates ground-truth observation from model inference, establishing immutable audit baselines.
S V

Refract →

Open sourceVerificationDeterministic

Open-source deterministic verification engine that checks whether a claim still holds against its sources. Produces structured change events with full provenance records.

Architecture
Provenance graph, source registration, claim-checking pipeline, structured event output. No model — pure observation.
Relevance
The infrastructure layer AI systems need to know when underlying facts change. Directly applicable to retrieval-augmented generation freshness monitoring.

Capability Graph →

Open sourceAgentsReliability

Prototype for tracking what an AI agent can do, what may be decaying, and the dependencies behind new capabilities. Built with Claude Code.

Architecture
Maturity scores, decay rates, bottleneck analysis, unlock paths. Directed graph of capabilities with edge weights for cost of acquisition.
Production application
Provides directed graph inspection to detect agent capability drift and dependency decay in production.

Operational evidence

Shipped production systems across healthcare — a domain where the cost of automated error is measured in patient outcomes.

Epic EHR — Implementation Engineer

EHRDecision supportPolicy

Configured clinical decision support and quality-measurement workflows within Epic EHR for federal incentive programs. Mapped institutional policies into system rules.

Relevance to AI
Understood first-hand how alert fatigue, override rates, and institutional trust interact when automated recommendations meet clinical judgment.
Scope
Implemented decision-support and quality-measurement workflows within Epic's vendor-supplied clinical rule engine.
Dr. Smith's Office

Doximity Dialer — Product Lead

CommunicationRegulatedScaleHIPAA

Product lead for Doximity Dialer — a regulated clinician communication product. 110M+ calls from 300K+ active clinicians. Epic Haiku integration, HIPAA compliance.

Relevance to AI
Shipped a product where compliance, reliability, and user trust were non-negotiable. Managed the gap between what the product promised and what institutions would adopt.
Outcome
110M+ calls, 4.8-star iOS rating (up from 3.7).

Transcarent — Director of Product

Care navigationUtilizationOperations

Led product across 4 specialty-care programs (Surgery, Urgent Care, Behavioral Health, Oncology). Designed utilization rules that reduced avoidable escalations while keeping high-risk cases clinician-led.

How this informs AI work
This experience informs my approach to rule-based routing, human review, and fallback design for AI-assisted workflows.
Outcome
Completed care plans as success metric (not engagement). Unified cross-program member record for consistent triage.

CancerCompass / CTCA — Director of Digital Products

OncologyContentContent moderation

Led product strategy and $2M execution for an oncology platform serving 30MM annual visitors. Built feedback loops with nursing teams to tune triage rules and reduce escalations.

How this informs AI work
Clinical content review and nursing feedback loops inform my approach to review gates for AI-assisted content systems.
Outcome
25% bounce rate reduction, 267% chat conversion increase. Platform later acquired by City of Hope.

Andwise — Co-Founder and CEO

FounderRegulatedRecommendation

Shipped regulated recommendation and review workflows for physician financial decisions. Combined legal oversight (contract analysis) with community-vetted provider directories.

Analogous AI deployment problem
AI-assisted guidance similarly needs explicit scope, qualified review, and an accountable escalation path.
Outcome
$240K raised, 1,200+ physician users, 50+ medical advisory board.