RLHF Specialist – Remote
Odixcity Consulting
Description du poste
About the role
We are looking for an RLHF Specialist to improve and align AI language models using Reinforcement Learning from Human Feedback. The role is fully remote, allowing you to collaborate with a global team to design and optimise feedback pipelines that enhance model safety, factual accuracy and alignment with human values.
Key responsibilities
- Generate high-quality preference data by comparing model responses and ranking them on helpfulness, honesty and harmlessness.
- Design multi-turn prompts to stress-test model behaviour and expose reasoning or safety weaknesses.
- Write detailed chain-of-thought explanations to train reward models.
- Collaborate with ML engineers to analyse failure modes, identify data gaps and propose data-driven interventions.
- Develop and iterate annotation strategies, ensuring consistency across a distributed team.
- Probe models for vulnerabilities, biases or hallucinations and document findings.
- Maintain a personal test set of prompts to monitor performance over time and re-evaluate new model versions against historical benchmarks.
- Translate complex RL concepts into repeatable tasks for junior annotators and create templated instruction sets.
Required profile
- Minimum 2 years experience in data annotation, model evaluation, computational linguistics or trust-and-safety for AI/ML.
- Strong proficiency in Python and at least one deep-learning framework (PyTorch, JAX or TensorFlow).
- Deep understanding of reinforcement-learning concepts such as PPO, trust-regions and reward hacking.
- Hands‑on experience fine‑tuning open‑source LLMs (e.g., Llama 2/3, Mistral, Gemma) using LoRA/QLoRA.
- Experience with annotation tools (LabelBox, Scale AI, Snorkel) and human-in-the-loop workflows.
- Familiarity with constitutional AI, self-alignment techniques and open-source alignment libraries.
- Experience with cloud platforms (AWS SageMaker or GCP Vertex AI).
Required skills
- Python
- PyTorch
- JAX
- TensorFlow
- LoRA/QLoRA fine-tuning
- LabelBox
- Scale AI
- Snorkel
- AWS SageMaker
- GCP Vertex AI
Questions fréquentes
Pourquoi signalez-vous cette offre ?
Aller plus loin
Salaires, guides et recherches au Gabon.
Postulez en 30 secondes
Entrez votre email pour postuler. Un compte sera cree automatiquement.
En continuant, vous acceptez nos conditions d'utilisation.
Deja un compte ? Connexion
Publie il y a 1 mois
Expire dans 5 jours
94 vues · 0 interesses
Boostez vos chances
Importez votre CV : nous vous proposons les offres qui matchent votre profil.
Analyse de votre CV en cours...
Odixcity Consulting
Offres similaires
-
Data Annotator (French) – Freelance
Innodata Inc. Gabon -
Dataset Curator – AI & ML Data Specialist
Odixcity Consulting Gabon -
Quality Assurance Analyst
Odixcity Consulting Gabon -
Développeur Go (Golang) – microservices backend
RED TIC Libreville -
Développeur Cloud / Back‑end .NET Azure (Freelance)
RED TIC Libreville