Celebrating 30 Years of Idealist! Discover where we’ve been, and where we’re going.
Nonprofit
Published 15 days ago

Volunteer AI Engineer (Evaluations & Benchmarking) — AI Apps for Human Communication

Remote, Volunteer can be anywhere in the world
I Want to Help


  • Details

    Time Commitment:
    Flexible
    Commitment Details:
    We hope volunteers can contribute approximately 8 or more hours per week, although hours are flexible.
    Recurrence:
    Recurring
    Volunteers Needed:
    2
    Cause Areas:
    Community Development, Education, Human Rights & Civil Liberties, Policy

    Description

    Volunteer AI Engineer (Evaluations & Benchmarking) — AI Apps for Human Communication

    Help build the rigorous evaluation pipelines and benchmarks that ensure conversational AI agents remain objective, depolarizing, emotionally safe, and constructive.

    Crossing Party Lines is seeking a detail-oriented AI Engineer specializing in Evaluations (Evals) to join our Apps and Agents team. Evaluating conversational agents designed for politically charged, high-stakes dialogue requires moving beyond basic accuracy checks. You will define, implement, and automate evaluation suites that measure nuanced qualities such as neutrality, tone calibration, reframing effectiveness, cognitive empathy, and resistance to bias or sycophancy.

    About the Apps and Agents Team

    Our team blends generative AI models and multi-agent workflows with Crossing Party Lines’ proprietary depolarization frameworks. As we prepare applications for Y Combinator and work toward broader public deployment, reliable evaluation is central to building tools that users and communities can trust.

    What You’ll Do

    • Design and maintain systematic evaluation datasets (golden sets) reflecting diverse, realistic conversational conflict and political viewpoints.
    • Implement multi-dimensional evaluation rubrics—combining deterministic checks, reference-based metrics, and LLM-as-a-judge pipelines.
    • Benchmark agent outputs against specific communication standards: active listening, cognitive perspective-taking, de-escalation, and non-judgmental feedback.
    • Run regression tests, track performance across prompt and model revisions, and integrate automated eval gates into CI/CD workflows.
    • Conduct adversarial red-teaming to uncover subtle bias, conversational failure modes, hallucinations, or toxic drift.
    • Collaborate with depolarization experts, product designers, and backend developers to calibrate evaluation criteria against real-world human judgment.

    You May Be a Good Fit If You

    • Have experience with Python and LLM application frameworks (e.g., LangChain/LangGraph, LlamaIndex, or raw API tool orchestration).
    • Understand modern LLM evaluation concepts and tools (e.g., promptfoo, DeepEval, Ragas, Phoenix, LangSmith, or custom LLM-as-a-judge scorers).
    • Have a keen analytical mindset and can translate qualitative human communication principles into clear quantitative metrics.
    • Are comfortable analyzing edge cases across a wide spectrum of social and political perspectives without imposing personal viewpoints.
    • Value transparency, scientific rigor, and reproducibility in AI benchmarking.

    What You’ll Gain

    • Direct experience solving one of the most critical challenges in production AI: behavioral evaluation and alignment for multi-turn conversational agents.
    • A key engineering role on early-stage AI products being prepped for high-profile accelerators and widespread adoption.
    • Collaboration with communication experts, prompt engineers, and product builders.
    • Professional references recognizing your technical rigor and contributions to production-grade AI reliability.

    Time, location, and support

    This is a remote volunteer position. You may complete most of the work on your own schedule.

    We hope volunteers can contribute approximately 8 or more hours per week, although hours are flexible. Team members should also be available for one or two online team meetings per week.

    You will report to Lisa Swallow, co-founder of Crossing Party Lines and lead of the Apps and Agents team. You will work closely with other volunteers and developers and will receive background information, product specifications, and access to the CPL frameworks underlying the applications.

    We welcome applicants from different backgrounds, professions, political viewpoints, and lived experiences. Crossing Party Lines does not ask volunteers to agree politically. We ask them to approach differences with curiosity, respect, and a sincere desire to understand.

    Location

    Remote
    Volunteer can be anywhere in the world
    Associated Location
    244 5th Ave suite l208, New York, NY 10001, USA

    Please fill out this form

    Instructions:

    Please send us a brief introduction explaining:

    • Why this work interests you
    • Your AI evaluation and benchmarking experience
    • Approximately how much time you are available each week
    • A link to a portfolio or examples of relevant work, when available
    All fields are required
    I acknowledge that use of the Idealist Applicant Tracking System is subject to Idealist's Privacy Policy and Terms of Service.
    Illustration

    Discover Your Calling

    Find opportunities to change the world with the latest social-impact job, internship, and volunteer listings. Plus, explore resources for taking action in your community.