📢 Nouveau : recevez les offres du jour sur notre canal WhatsApp
Jobiglo

Aucun resultat.

RLHF Specialist – Remote

Odixcity Consulting

Remote
Remote Mid 🇬🇧 English
Python PyTorch JAX TensorFlow LabelBox Scale AI Snorkel AWS SageMaker GCP Vertex AI

Description du poste

About the role

We are looking for an RLHF Specialist to improve and align AI language models using Reinforcement Learning from Human Feedback. The role is fully remote, allowing you to collaborate with a global team to design and optimise feedback pipelines that enhance model safety, factual accuracy and alignment with human values.

Key responsibilities

  • Generate high-quality preference data by comparing model responses and ranking them on helpfulness, honesty and harmlessness.
  • Design multi-turn prompts to stress-test model behaviour and expose reasoning or safety weaknesses.
  • Write detailed chain-of-thought explanations to train reward models.
  • Collaborate with ML engineers to analyse failure modes, identify data gaps and propose data-driven interventions.
  • Develop and iterate annotation strategies, ensuring consistency across a distributed team.
  • Probe models for vulnerabilities, biases or hallucinations and document findings.
  • Maintain a personal test set of prompts to monitor performance over time and re-evaluate new model versions against historical benchmarks.
  • Translate complex RL concepts into repeatable tasks for junior annotators and create templated instruction sets.

Required profile

  • Minimum 2 years experience in data annotation, model evaluation, computational linguistics or trust-and-safety for AI/ML.
  • Strong proficiency in Python and at least one deep-learning framework (PyTorch, JAX or TensorFlow).
  • Deep understanding of reinforcement-learning concepts such as PPO, trust-regions and reward hacking.
  • Hands‑on experience fine‑tuning open‑source LLMs (e.g., Llama 2/3, Mistral, Gemma) using LoRA/QLoRA.
  • Experience with annotation tools (LabelBox, Scale AI, Snorkel) and human-in-the-loop workflows.
  • Familiarity with constitutional AI, self-alignment techniques and open-source alignment libraries.
  • Experience with cloud platforms (AWS SageMaker or GCP Vertex AI).

Required skills

  • Python
  • PyTorch
  • JAX
  • TensorFlow
  • LoRA/QLoRA fine-tuning
  • LabelBox
  • Scale AI
  • Snorkel
  • AWS SageMaker
  • GCP Vertex AI

Questions fréquentes

Le salaire n'est pas communiqué publiquement par le recruteur. Vous pouvez postuler et négocier directement avec Odixcity Consulting.
Cliquez sur "Postuler maintenant" en haut de la page. Vous pouvez importer votre CV en 1 clic — Jobiglo extrait automatiquement vos informations et postule pour vous.

Pourquoi signalez-vous cette offre ?

Merci pour votre signalement. Nous allons examiner cette offre.

Postulez en 30 secondes

Entrez votre email pour postuler. Un compte sera cree automatiquement.

En continuant, vous acceptez nos conditions d'utilisation.

Deja un compte ? Connexion

💬 Contactez-nous sur Telegram Discuter sur WhatsApp

Publie il y a 1 mois

Expire dans 5 jours

94 vues · 0 interesses

Boostez vos chances

Importez votre CV : nous vous proposons les offres qui matchent votre profil.

Analyse de votre CV en cours...

Odixcity Consulting