Jobiglo

Sem resultados

RLHF Specialist

Odixcity Consulting

Remote
Remote Mid 🇬🇧 English
Python PyTorch JAX TensorFlow PPO Trust Regions Reward Hacking LoRA QLoRA Llama 2 Mistral Gemma LabelBox Scale AI Snorkel AWS SageMaker GCP Vertex AI

Descrição do cargo

About the role

The RLHF Specialist will help improve large language models by designing and operating Reinforcement Learning from Human Feedback pipelines. Working remotely with a global team, you will ensure that AI outputs are helpful, honest, and harmless.

Key responsibilities

  • Generate high‑quality preference data by ranking model responses on helpfulness, honesty, and harmlessness.
  • Design multi‑turn prompts to stress‑test model reasoning and safety.
  • Write chain‑of‑thought explanations to train reward models.
  • Collaborate with ML engineers to analyse failure modes and close data gaps.
  • Develop and iterate annotation strategies for consistent preference scoring.
  • Probe models for biases, hallucinations, and vulnerabilities, documenting findings.
  • Analyze edge cases where reward models behave unexpectedly and suggest data interventions.
  • Create templated instruction sets for large annotation teams and translate RL concepts into repeatable tasks.
  • Maintain a personal test set of prompts to monitor model performance over time.

Required profile

  • At least 2 years of experience in data annotation, model evaluation, computational linguistics, or AI trust & safety.
  • Strong proficiency in Python and deep‑learning frameworks (PyTorch, JAX, or TensorFlow).
  • Deep understanding of reinforcement‑learning concepts such as PPO, trust regions, and reward hacking.
  • Hands‑on experience fine‑tuning open‑source models (e.g., Llama 2/3, Mistral, Gemma) using LoRA/QLoRA.
  • Experience with annotation platforms (LabelBox, Scale AI, Snorkel) and human‑in‑the‑loop workflows.
  • Familiarity with cloud AI services (AWS SageMaker, GCP Vertex AI).

Required skills

  • Python
  • PyTorch
  • JAX
  • TensorFlow
  • PPO, Trust Regions, Reward Hacking
  • LoRA / QLoRA
  • Llama 2, Llama 3, Mistral, Gemma
  • LabelBox, Scale AI, Snorkel
  • AWS SageMaker
  • GCP Vertex AI

Questions fréquentes

Le salaire n'est pas communiqué publiquement par le recruteur. Vous pouvez postuler et négocier directement avec Odixcity Consulting.
Cliquez sur "Postuler maintenant" en haut de la page. Vous pouvez importer votre CV en 1 clic — Jobiglo extrait automatiquement vos informations et postule pour vous.

Motivo do reporte

Obrigado! A sua denúncia foi enviada aos administradores.

Candidate‑se em 30 segundos

Introduza o seu e‑mail para candidatar‑se. Uma conta será criada automaticamente.

Ao continuar, aceita os nossos termos de uso.

Já tem uma conta? Entrar

💬 Chat with us on Telegram Conversar no WhatsApp

Publicado há 1 mês

Expira em 1 dia

46 visualizações · 0 interested

Aumente suas chances

Envie seu CV: vamos sugerir as vagas que combinam com seu perfil.

A analisar o seu CV...

Odixcity Consulting