¡Celebramos 30 años de Idealist! Descubre nuestra trayectoria y nuestra visión para el futuro.
Organización sin fin de lucro
Publicado hace 16 días

Volunteer AI Engineer (Evaluations & Benchmarking) — AI Apps for Human Communication

A distancia, El/la voluntario/a puede estar en cualquier país del mundo
Quiero ayudar


  • Descripción

    Flexibilidad:
    Flexible
    Detalles del compromiso:
    We hope volunteers can contribute approximately 8 or more hours per week, although hours are flexible.
    Frecuencia:
    Recurrente
    Buscando personas voluntarias:
    2
    Área de impacto:
    Desarrollo de comunidades, Educación, Derechos humanos y libertades civiles, Política

    Descripción

    Volunteer AI Engineer (Evaluations & Benchmarking) — AI Apps for Human Communication

    Help build the rigorous evaluation pipelines and benchmarks that ensure conversational AI agents remain objective, depolarizing, emotionally safe, and constructive.

    Crossing Party Lines is seeking a detail-oriented AI Engineer specializing in Evaluations (Evals) to join our Apps and Agents team. Evaluating conversational agents designed for politically charged, high-stakes dialogue requires moving beyond basic accuracy checks. You will define, implement, and automate evaluation suites that measure nuanced qualities such as neutrality, tone calibration, reframing effectiveness, cognitive empathy, and resistance to bias or sycophancy.

    About the Apps and Agents Team

    Our team blends generative AI models and multi-agent workflows with Crossing Party Lines’ proprietary depolarization frameworks. As we prepare applications for Y Combinator and work toward broader public deployment, reliable evaluation is central to building tools that users and communities can trust.

    What You’ll Do

    • Design and maintain systematic evaluation datasets (golden sets) reflecting diverse, realistic conversational conflict and political viewpoints.
    • Implement multi-dimensional evaluation rubrics—combining deterministic checks, reference-based metrics, and LLM-as-a-judge pipelines.
    • Benchmark agent outputs against specific communication standards: active listening, cognitive perspective-taking, de-escalation, and non-judgmental feedback.
    • Run regression tests, track performance across prompt and model revisions, and integrate automated eval gates into CI/CD workflows.
    • Conduct adversarial red-teaming to uncover subtle bias, conversational failure modes, hallucinations, or toxic drift.
    • Collaborate with depolarization experts, product designers, and backend developers to calibrate evaluation criteria against real-world human judgment.

    You May Be a Good Fit If You

    • Have experience with Python and LLM application frameworks (e.g., LangChain/LangGraph, LlamaIndex, or raw API tool orchestration).
    • Understand modern LLM evaluation concepts and tools (e.g., promptfoo, DeepEval, Ragas, Phoenix, LangSmith, or custom LLM-as-a-judge scorers).
    • Have a keen analytical mindset and can translate qualitative human communication principles into clear quantitative metrics.
    • Are comfortable analyzing edge cases across a wide spectrum of social and political perspectives without imposing personal viewpoints.
    • Value transparency, scientific rigor, and reproducibility in AI benchmarking.

    What You’ll Gain

    • Direct experience solving one of the most critical challenges in production AI: behavioral evaluation and alignment for multi-turn conversational agents.
    • A key engineering role on early-stage AI products being prepped for high-profile accelerators and widespread adoption.
    • Collaboration with communication experts, prompt engineers, and product builders.
    • Professional references recognizing your technical rigor and contributions to production-grade AI reliability.

    Time, location, and support

    This is a remote volunteer position. You may complete most of the work on your own schedule.

    We hope volunteers can contribute approximately 8 or more hours per week, although hours are flexible. Team members should also be available for one or two online team meetings per week.

    You will report to Lisa Swallow, co-founder of Crossing Party Lines and lead of the Apps and Agents team. You will work closely with other volunteers and developers and will receive background information, product specifications, and access to the CPL frameworks underlying the applications.

    We welcome applicants from different backgrounds, professions, political viewpoints, and lived experiences. Crossing Party Lines does not ask volunteers to agree politically. We ask them to approach differences with curiosity, respect, and a sincere desire to understand.

    Ubicación

    A distancia
    La persona voluntaria puede estar en cualquier lugar del mundo
    Ubicación asociada
    244 5th Ave suite l208, New York, NY 10001, USA

    Por favor, llena este formulario

    Instrucciones:

    Please send us a brief introduction explaining:

    • Why this work interests you
    • Your AI evaluation and benchmarking experience
    • Approximately how much time you are available each week
    • A link to a portfolio or examples of relevant work, when available
    Todos los campos son obligatorios
    Entiendo que el uso de la herramienta de seguimiento de candidaturas de Idealist está sujeto a la Política de Privacidad de Idealist y a los Términos del Servicio.
    Illustration

    Descubre tu vocación

    Encuentra oportunidades para cambiar el mundo con las últimas oportunidades de empleo, pasantías/prácticas y voluntariado con impacto social. Además, podrás explorar recursos para generar impacto positivo en tu comunidad.