How can you make a difference?
- Prepare, clean and structure the agreed test prompt set.
- Conduct extensive testing of agreed prompts across multiple LLM providers and configurations, and curate the most informative model responses for expert review.
- Run evaluations under different approved safety/content-filter configurations and document how these settings affect model behaviour.
- Capture model responses and relevant test metadata.
- Troubleshoot API, model, safety-filter or platform issues.
- Develop and maintain code to call commercial LLM APIs and, where required, deploy and interact with open-source models for evaluation.
- Conduct targeted reviews of state-of-the-art child-safety and AI-safety research to identify relevant prompts, failure modes and emerging risks.
- Organise outputs into a simple format for expert review.
- Support simple rating, ranking and reviewer comments.
- Capture review results in a structured, exportable format.
- Iterate the prototype based on expert feedback.
- Participate in communication with reviewers, collect feedback on both model behaviour and the review process, and refine the testing/review infrastructure accordingly.
- Support re-testing when experts identify new failure modes, ambiguities or edge cases.
- Help translate expert feedback into new or refined test cases.
- Maintain a clean, versioned prompt-response dataset.
Minimum requirements
data science, machine learning, information security, human-computer interaction or a related technical
field. A first university degree combined with additional relevant professional experience may be accepted
in lieu of an advanced degree, subject to UNICEF requirements.
Enter Disciplines:
Computer Science, Artificial Intelligence, Data Science, Machine Learning, Information Security or related
field.
data science, model evaluation or a closely related technical field.
Demonstrated hands-on experience evaluating LLM outputs, including qualitative assessment across
multiple models, versions or configurations.
Practical experience with prompt testing, AI safety testing and/or AI red teaming.
Strong Python skills and experience working with LLM APIs and
Ability to design and execute structured, repeatable tests and capture sufficient metadata for
reproducibility.
Understanding of system prompts, content filters, moderation layers, refusal behaviour and other safety
controls that may affect evaluation results.
Proficiency working with CSV, JSON and JSONL and preparing clear prompt-response datasets for expert
review.
Experience troubleshooting model, API, filtering or evaluation-platform issues.
Experience deploying and running open-source LLMs in cloud environments, including model serving and
inference workflows.
Ability to communicate technical findings clearly and work iteratively with non-technical subject-matter experts.
- Language Requirements: Fluency in English is required Desirables: Azure AI Foundry or comparable evaluation-platform experience; experience in child safety, online harms, safeguarding or responsible AI; experience prototyping lightweight review interfaces.
For every Child, you demonstrate...
UNICEF does not hire candidates who are married to children (persons under 18). UNICEF has a zero-tolerance policy on conduct that is incompatible with the aims and objectives of the United Nations and UNICEF, including sexual exploitation and abuse, sexual harassment, abuse of authority and discrimination based on gender, nationality, age, race, sexual orientation, religious or ethnic background or disabilities. UNICEF is committed to promote the protection and safeguarding of all children. All selected candidates will, therefore, undergo rigorous reference and background checks, and will be expected to adhere to these standards and principles. Background checks will include the verification of academic credential(s) and employment history. Selected candidates may be required to provide additional information to conduct a background check, and selected candidates with disabilities may be requested to submit supporting documentation in relation to their disability confidentially.
- An up-to-date TMS profile and curriculum vitae (CV)
- Cover letter
UNICEF does not charge a processing fee at any stage of its recruitment, selection, and hiring processes (i.e., application stage, interview stage, validation stage, or appointment and training). UNICEF will not ask for applicants’ bank account information.
All UNICEF positions are advertised, and only shortlisted candidates will be contacted and advance to the next stage of the selection process.