Joseph DeMasse · Utica, Michigan

AI / ML engineer.
Whole-team output, one person.

At Ur Mortgage I run the entire engineering function — lead developer, project manager, AI/ML engineer and full-stack developer in one seat. I design, train, ship and operate production AI end-to-end: autonomous voice agents on the public telephone network, self-training retrieval brains with compliance firewalls, from-scratch machine learning, speech DSP in Rust, AVX2 kernels in Go — and the Salesforce + AWS platform underneath. In a federally regulated industry, running every day.

See the systems Download resume (PDF)
real calls mined into a training corpus
0
calls in one 24-hour window, 37 concurrent
0
languages shipped in production code
0
tests green on the underwriting engine
0

Every number on this page is real and traceable to a running system.

The systems

Seven flagship platforms — each designed, built, deployed and operated solo — plus the fleet behind them. Four named AI models run across them: Avery and Alexa on the phones, Lilly in training, Mario in underwriting.

AOPE — the autonomous AI dialer

● in production · Node.js · Telnyx · Twilio · OpenAI Realtime · Electron · AWS

A Five9-class outbound calling platform where AI voice agents — Avery and Alexa — place the calls, hold natural full-duplex conversations with borrowers, capture 1003 mortgage-application fields into Salesforce live during the call, and warm-transfer qualified, interested borrowers to a licensed loan officer. Runs a 22-user floor on an always-on AWS box; peaked at 25,864 calls in 24 hours across 37 concurrent legs, fed by a 672,705-lead Salesforce pool.

  • Real-time voice engineering — sub-second barge-in, adaptive jitter buffers, μ-law G.711 media, per-agent WebRTC stations, carrier-side Telnyx conferences with hold / consult / merge / supervisor-listen as membership operations, and a one-button warm-transfer flow (hold → find LO → private brief → bridge → drop off).
  • Five commercial voice stacks integrated — OpenAI Realtime, ElevenLabs, Cartesia, Hume and NVIDIA PersonaPlex — with cloned voices, per-call TTS arms and a scheduled four-engine bakeoff harness that graded them on live leads.
  • In-house ML engines, $0 per call — an empirical-Bayes connect-likelihood model (trained on 4,538 calls across 220 time-slots) that decides who to dial and when, a compliance-NLP engine that scans live transcripts and auto-suppresses opt-outs to DNC, best-time-to-call prediction, ANI-rotation bandits over a 100-number pool, and a nightly self-retrain cycle.
  • Speech DSP down to the sample — a Rust prosody analyzer (mel-spectrogram front-end → streaming GRU multi-task heads for emotion, voicemail, IVR and AI-screener detection over raw telephony audio) and a pure-Node acoustic answering-machine detector built on Goertzel tone analysis (~150 ms beep detection).
  • Compliance as code, fail-closed — every dial clears DNC (internal + national), TCPA calling windows (timezone/DST-aware), state licensing, consent and per-state recording-disclosure gates before it rings; frequency caps; hash-chained forensic call logs; RBAC-locked admin console; Ed25519-signed auto-updating Electron agent stations.
  • Operated like a product — 20-module admin console, tiered skill routing with drag-and-drop, campaign engine with cadence retries, failure feed with automated diagnosis, disaster-recovery repo with a from-zero restore playbook.

Astravyx — the AI brain

● in production · Python · Haystack · ChromaDB · Ollama · numpy · PyTorch · Go · Rust

A self-training retrieval-augmented intelligence platform organised like a brain — seven cortex “lobes” (retrieval hippocampus, classifier neocortex, auditory perception, scoring basal ganglia, judgment prefrontal, language, motor/tools) over an 11,000+-chunk domain-tagged knowledge index, with a hard rule at its core: the judgment lobe is never neural, and no mortgage fact is ever taught that isn’t grounded in cited guideline text.

  • Hybrid retrieval — vector search (nomic-embed-text embeddings in ChromaDB) fused with BM25 and an ONNX cross-encoder reranker; per-chunk domain metadata; content-hash dedup; corruption-safe online backups.
  • The dojo — a self-improving training loop — clean → augment → quiz → teach → re-quiz, fully unattended in daemon mode. Cloud models generate exams, the local model answers, a judge grades, and failed items are taught back as retrievable gold cards. Book-grounded augmentation doubled a benchmark quiz from 33% to 67% in one round — the “correct data, not more data” lever, proven.
  • Compliance firewall inside the training loop — a gold answer is taught only if every number in it appears in retrieved guideline text; verified to accept a real “DTI 50%” and refuse a fabricated “DTI 73%”. Fail-closed: retrieval error → not grounded → not taught. ECOA / Reg-B by construction.
  • Own-model programme (Forge) — LoRA/QLoRA fine-tune pipeline over the transformers/PEFT stack with dataset export, adapter merge, A/B, promote and rollback stages — every promotion gated by a champion/challenger battery over hand-grounded golden sets (citation-grounding floor 0.95, false-refusal creep ceiling, stale-baseline invalidation by design).
  • From-scratch ML layer — pure-numpy MLP document classifier (768→128→softmax, 0.89 test accuracy vs 0.68 baseline), K-Means + silhouette clustering, PCA embedding maps, distance + kNN-density anomaly detection, and self-contained offline HTML visualizers with animated learning curves — no sklearn, no torch on the hot path.
  • Apex lead engine — a from-scratch second-order GBDT ensembled with logistic regression, isotonic calibration (ECE 0.10 → 0.029), split-conformal confidence, expected-value ranking and two-model causal uplift; validated end-to-end on synthetic cohorts (AUC 0.72, top-decile EV-lift 2.47×) before touching real funnel data.
  • Custodian, Sentinel & the immune system — an isolation-forest data-quality steward that audits the brain against 17 counterpart systems (it found and reversibly quarantined ~44% index pollution), plus UEBA/DLP-style monitoring — with the learned rule that a statistical anomaly may flag, but only definitive evidence may delete.
  • Speaks every protocol — FastAPI desktop app with 3D brain visualization, 29 MCP tools for agent interop, REST APIs, and Go/Rust sidecar services where the CPU work lives.

Crucible + Mario — the AI underwriting processor

● in production · Python · Go + AVX2 · Rust · Salesforce · FastAPI · MeshyAI

An adversarial, transparent AI underwriting engine — advocate, adversary and adjudicator argue every loan into a structured path to yes — wired into a document-intelligence processor that reads what arrives and reasons about what it proves. ~52 modules, 449 tests green. It computes, cites and recommends; a human owns every credit decision, and one ECOA/Reg-B firewall gates every human-facing sentence.

  • Document intelligence — 53 document kinds × 4 signal classes (including disqualifiers), 93 evidence-span extractors, condition-intent mining from the English text itself (survives investor renumbering), document-set reconciliation and LOX drafting.
  • Real underwriting math — self-employed income add-backs, declining-income guards that never average a fall upward, asset haircuts and unsourced-deposit deductions, DTI at the note rate with buydowns flagged (never used), residual income, LLPA/MI pricing cliffs with the clean levers to step over them, and 24 remediation pathways.
  • Decision science — counterfactual close-probability (“where an hour of work pays”), a UCB contextual bandit that learns which action helps per situation, and AUS integration (DU/LP) with reconciliation.
  • Performance engineering — Go batch-scoring kernels with hand-written AVX2 assembly, differentially tested against a portable oracle: 8.7× dot-product, 9.0× batch scoring — 100,000 loans in 214 µs; a Rust PDF parser (lopdf) rebuilt to 100% text recall after the Go original plateaued at 47%.
  • Mario — the face of the engine — an AI underwriting assistant living directly on the Salesforce loan record as an animated 3D processor-bot: a real rigged, walking 3D avatar (MeshyAI image-to-3D → auto-rig → GLB, rendered via model-viewer) with a chat overlay. Ask Mario “what needs to be done” and he lists the actual open conditions and why each is open, straight from the engine.
  • Grounded chat with a self-critic — Mario’s answers pass a refute-before-answer critic and abstain rather than hallucinate; he serves the graded underwriting memo, a durable audit log (custom Salesforce object), and a rung-gated action bar. Deployed to production and sandbox orgs.
  • Delivered as real infrastructure — a hosted FastAPI service behind a Caddy reverse proxy that injects the API key server-side (Salesforce never holds the secret), reached via Named Credential; REST + MCP surfaces; and a governed autonomy dial: observe → stamp → outreach → submit.

Lilly — an own-weights speech-to-speech model

architecture complete · training staged · PyTorch · Rust · neural codecs

Astravyx’s sovereign voice: a full-duplex speech-to-speech model designed from the research up — own neural codec, own language model, own weights — so no external vendor sits inside the phone call.

  • LillyCodec — SEANet encoder/decoder with 16-codebook residual vector quantization (codebook-0 semantic), 24 kHz audio at a 12.5 Hz token rate.
  • LillyS2SLM — an RQ-Transformer: a 7B-class temporal backbone over time plus a depth transformer across codebooks, dual-stream (caller + Lilly modeled jointly so turn-taking is learned), with an inner-monologue text head and acoustic delay patterns for streaming.
  • The data moat — a Whisper + diarization pipeline over 203,606 real recorded sales calls, PII-scrubbed and aligned into dual-stream training shards.
  • Honest engineering — the serving bridge, eval-parity harness against a commercial baseline, and phase-gated training plan (codec → LM → promote) are built; weight training is staged pending dedicated GPUs, and the README says exactly that.

The Salesforce ↔ UWM lending platform

● in production · Apex · LWC · AWS Lambda · DynamoDB · MISMO XML · GraphQL

The integration backbone of a live mortgage operation: three Salesforce orgs (prod / dev / QA) wired to America’s largest wholesale lender through an AWS serverless layer — used by the whole company, every day.

  • Loan lifecycle automation — MISMO 3.4 XML loan creation, date-tracking webhooks that drive Opportunity stages, conditions sync with a command-center LWC, loan export, AUS + decoupled credit-pull LWCs, and HELOC routing fixes negotiated down to XML element placement with the lender’s own engineers.
  • Reliability engineering in Apex — webhook race conditions solved with deferred Queueable retries (plus a bulk replay script that healed 11 race-lost loans), idempotent upserts, validation-first multi-org deploys, and permission-set architecture across 30+ users per org.
  • Real-time lead flow — record-triggered Flows + invocable Apex push leads to the dialer within seconds of creation; the dialer writes status, attempts and 1003 fields back.
  • The web fleet — urmortgages.com (production cutover on AWS Amplify + Route 53 with an AI-generated daily learning-center article), a borrower client portal, a document hub with UWM-conditions cross-wiring, gamified team leaderboards, and an embedded React lead-command panel bridged into Salesforce via postMessage.

urmortgages.com — the AI-wired storefront

● in production · AWS Amplify · Route 53 · Lambda · SES · Gemini · JSON-LD

The company’s public site, taken to production on AWS end-to-end — Amplify hosting, a full Route 53 DNS migration that preserved live e-mail across six DKIM identities, security headers, and a content engine that publishes itself. Not a brochure: every page feeds the lead machine.

  • The OMA — online mortgage application — a multi-step application wizard with abandonment capture engineered in: the moment a visitor has left a name plus phone or e-mail, a five-minute idle timer (or tab close) upserts a hot Lead into Salesforce; finishing the application updates the same Lead — zero duplicates, via an external-id upsert key that survives privacy modes. TCPA consent is captured with timestamp and exact consent text, and every OMA lead lands priority-flagged for the dialer.
  • A Learning Center that writes itself — a daily scheduled pipeline where the AI stack authors one positive, mortgage-specific news article via Gemini with grounded web search (real cited sources), falls back to Astravyx/OpenAI topic rotation, runs a compliance scrub that neutralizes rate figures, renders full SEO, and updates the article grid, sitemap and RSS — unattended, every morning at 8.
  • AEO — built to be cited by AI — an answer-engine-optimization program targeting ChatGPT, Gemini, Claude and Grok: a FinancialService/WebSite JSON-LD entity graph, per-product Service schema, static NAP baked into raw HTML for non-JS AI crawlers, honest sitemap lastmod, and a Michigan product cluster whose FAQ schema mirrors the exact questions borrowers ask AI assistants.
  • The plumbing — Contact + Careers forms through an SES Lambda, 10DLC/TCR brand verification for compliant SMS, and a custom seo-inject build tool that bakes canonical tags, OG cards and schema into every page at deploy time.

The Five9 corpus excavation

completed · AWS S3 · faster-whisper · diarization · data forensics

Mined the company’s legacy call archive — ~984,000 objects in S3 — into an AI training asset: censused the bucket, filtered to 203,606 substantial recorded calls, and ran a multi-day local transcription + diarization pipeline (no cloud spend, no data egress) to build the corpus that teaches the voice agents how elite closers actually talk. The same forensics answered a business question the humans had argued about for months: the legacy dialer’s penetration edge was a 6,406-number caller-ID strategy — not cadence, not scripts — which set the roadmap for the in-house number-rotation engine.

…and the fleet

Sentinel AI — internal UEBA/DLP + dev-intelligence platform; autoencoder + K-Means per-actor anomaly detection.
Raquazya — intrusion / deploy-drift detection with a live dashboard.
Peacock — scope-gated, authorized pentest kill-chain runner (TypeScript).
JIRANAMO — signed auto-updating Electron dev-accountability dashboard (CWE-494 hardened).
God’s Eye View — 3D OSINT globe, locally deployed and extended.
Lati — 3D social / streaming platform prototype (Three.js).
Neural Forge — Next.js + Electron AI/ML study app built for the MIT programme.
Reliquary — digital-artifact exchange + geo-hunt app (React, Capacitor → iOS/Android).
LangHub — desktop toolchain manager that verifies 36 languages by compiling in each.
Agentopolis — 8-bit AI-agent life simulation (Pygame).
XMRFleet — RandomX fleet-mining control panel with GUI + embedded C.
Exodia / Lead Loader — throttled bulk lead pipelines with safety invariants (Tkinter GUI).
UR:Online — gamified sales-floor leaderboard scored from live Salesforce activity, with a 3D avatar cosmetics store (MeshyAI).
Client portal & Doc Hub — borrower-facing web apps wired to UWM conditions, deployed across three orgs on Amplify.

The arsenal

Every technology below is in a repository I wrote — not a keyword list. Rendered from the same data file as the PDF resume, so the two can never drift.

Experience

Education

MIT xPRO — Designing and Building AI Products and Services

Professional certificate · 6 CEUs · July 2026 · two Exemplary Assignment awards

Signed by Dimitris J. Bertsimas (Vice Provost for Open Learning, MIT Sloan) and Brian Subirana (former Director, MIT Auto-ID Lab). What made it different: the coursework wasn’t hypothetical — every deliverable was designed against systems I was already running in production.

  • The AI design process — the four-stage framework (Intelligence → Business Process → AI Technology → Tinkering), performance-metric definition, IP strategy, data-approach selection, and anticipating “AI cancers” before they metastasize.
  • The double-diamond capstone — discover / define / develop / deliver applied to Avery: a pilot scoped to one pod, one lead source, one state; a pre-committed 20% lift criterion against a matched human control; KPI reaction rules, drift monitoring and versioned rollback — designed explicitly to escape pilot purgatory and the prototype trap. Awarded Exemplary Assignment.
  • Human-computer interaction — designed the loan-officer console under Nielsen’s heuristics: pre-dial compliance gates as error prevention, live transcript + status chips as system visibility, 10-second undo as user control, chunked panels under the 7±2 rule, ~180 ms suggested-line latency.
  • GANs, held to account — a conditional image-to-image super-resolution GAN restoring borrower-submitted property/document photos, assessed for sustainability, feasibility and responsibility: bias testing across property types and neighborhoods, EXIF/PII stripping, “AI-enhanced” watermarking beside the untouched original, human review on every output. Awarded Exemplary Assignment.
  • Superminds & collective intelligence — designing organizational processes that combine human and machine cognition (sense / remember / decide / create / learn), applied to lead routing, coaching and R&D loops.
  • Responsible AI & ML-Ops governance — Microsoft’s six principles applied as engineering constraints; offline validation against accuracy and compliance thresholds before any model or policy deploys, with rollback to the last validated version.

Michigan State University

B.S. Computer Science · Full Stack Web Development Boot Camp (College of Engineering)

Computer science foundations plus the College of Engineering’s full-stack web development program — certificate signed by department chairperson Abdol-Hossein Esfahanian.

Credentials

Certificates, licenses and awards — each independently verifiable.

Contact

I’m extremely happy where I am — and 100% open to contract work and hard problems. Anything AI, voice, platform or integration shaped.