Measure whether a post-training change improved real conversation.
Mohamed A M Elansary, PhD — multimodel evaluation under uncertainty, conversational and agent evaluation sets, and quality-signal measurement for post-training loops.
Evaluation under uncertainty
- Six-plus years of multimodel, multi-basin forecast experiments across hydroclimates on Linux/HPC.
- Compared statistical and physically based stacks, quantified uncertainty, and reported regime-dependent failure modes rather than a single flattering score.
- That is the measurement analogue of identifying quality signals for real-world conversational performance.
Conversational agents
- Production GPT, Claude, and Gemini agent workflows with retrieval, routing, tenant isolation, provenance, and regression evaluation sets at Vertexium, including a multi-tenant conversational receptionist.
- That maps to inspecting multi-turn traces. It is not Character.AI consumer-chat product ownership or RLHF-at-scale.
- PhD (or equivalent) is a posted requirement; this profile includes an Environmental Engineering PhD.
Proposed first contribution
For one post-training, data, or serving change already in flight, define what performance in the real world means for conversational quality versus a score that is easy to move. Write a small failure taxonomy: metric movement without a conversational quality change, slice-specific collapse, engagement that hides incoherence, serving change that wins latency and loses quality. Stand up a small evaluation set with provenance on traces, compare simple baselines, attach uncertainty, and write a clear report before expanding alignment-loss, data-mix, or sampling work. This is a proposed measurement approach, not a claim of prior Character-internal work, RLHF-at-scale, or invented metrics.
Honest fit boundary
Character.AI's consumer-chat product stack and RLHF-at-scale are a stretch. I have not trained Character models, run RLHF or online DPO, designed Character-internal alignment losses or samplers, or claimed Character-internal work, and I do not invent metrics or safety research. The credible contribution is evaluation under uncertainty, conversational and agent evaluation harnesses, scientific/HPC rigor, and production data pipelines.
Role and location
Research Engineer, Post-Training (All Industry Levels) · Redwood City or New York City · Hybrid. Ashby lists workplaceType Hybrid and isRemote true. Onsite-day count is not published. Willing to relocate to Redwood City or New York City with a relocation package. Fully remote work is not asserted.
Posting compensation: “$225K – $400K • Offers Equity”. · Official role posting