📢 New: get today's jobs on our WhatsApp Channel
Jobiglo

No results.

This job is no longer available

This job expired on 25/09/2026. It no longer accepts applications.

RLHF Specialist – Remote

Odixcity Consulting

Remote
Remote Mid 🇬🇧 English
Python PyTorch JAX TensorFlow LabelBox Scale AI Snorkel AWS SageMaker GCP Vertex AI

Job description

About the role

We are looking for an RLHF Specialist to improve and align AI language models using Reinforcement Learning from Human Feedback. The role is fully remote, allowing you to collaborate with a global team to design and optimise feedback pipelines that enhance model safety, factual accuracy and alignment with human values.

Key responsibilities

  • Generate high-quality preference data by comparing model responses and ranking them on helpfulness, honesty and harmlessness.
  • Design multi-turn prompts to stress-test model behaviour and expose reasoning or safety weaknesses.
  • Write detailed chain-of-thought explanations to train reward models.
  • Collaborate with ML engineers to analyse failure modes, identify data gaps and propose data-driven interventions.
  • Develop and iterate annotation strategies, ensuring consistency across a distributed team.
  • Probe models for vulnerabilities, biases or hallucinations and document findings.
  • Maintain a personal test set of prompts to monitor performance over time and re-evaluate new model versions against historical benchmarks.
  • Translate complex RL concepts into repeatable tasks for junior annotators and create templated instruction sets.

Required profile

  • Minimum 2 years experience in data annotation, model evaluation, computational linguistics or trust-and-safety for AI/ML.
  • Strong proficiency in Python and at least one deep-learning framework (PyTorch, JAX or TensorFlow).
  • Deep understanding of reinforcement-learning concepts such as PPO, trust-regions and reward hacking.
  • Hands‑on experience fine‑tuning open‑source LLMs (e.g., Llama 2/3, Mistral, Gemma) using LoRA/QLoRA.
  • Experience with annotation tools (LabelBox, Scale AI, Snorkel) and human-in-the-loop workflows.
  • Familiarity with constitutional AI, self-alignment techniques and open-source alignment libraries.
  • Experience with cloud platforms (AWS SageMaker or GCP Vertex AI).

Required skills

  • Python
  • PyTorch
  • JAX
  • TensorFlow
  • LoRA/QLoRA fine-tuning
  • LabelBox
  • Scale AI
  • Snorkel
  • AWS SageMaker
  • GCP Vertex AI

Questions fréquentes

Le salaire n'est pas communiqué publiquement par le recruteur. Vous pouvez postuler et négocier directement avec Odixcity Consulting.
Cliquez sur "Postuler maintenant" en haut de la page. Vous pouvez importer votre CV en 1 clic — Jobiglo extrait automatiquement vos informations et postule pour vous.

Why are you reporting this job?

Thank you for your report. We will review this job.

A question about this job?

Ask it here: you will get the full job summary by e-mail, right away.

💬 Chat with us on Telegram

Published 2 months ago

103 views · 0 interested

Boost your chances

Upload your CV — we will match you with relevant openings.

Analyzing your CV...

Odixcity Consulting