Applied Scientist (LLM)
⚲ Wrocław, Brno
Do uzgodnienia
Wymagania
- Machine Learning
- LLM
- Python
Opis stanowiska
Team Summary
Our distributed team is looking for an experienced Applied Scientist with a strong background in Large Language models to develop high-performance Generative AI features across Cloud and Edge environments.
Job Summary
In this role you will drive the transition from research to production by optimizing local inference through model compression and quantization for private, real-time Edge performance, while also engineering scalable RAG architectures and multi-agent systems for Cloud deployment. Your daily responsibilities encompass the full research lifecycle, including formulating hypotheses, generating synthetic datasets, fine-tuning LLMs, and validating safety and alignment, ultimately culminating in technical reports.
Responsibilities and Duties
• Design and implement advanced methods in prompt orchestration, fine-tuning (SFT/RLHF/DPO), and autonomous agentic workflows
• Curate high-quality training data from large-scale text and multi-modal sources
• Identify patterns in model hallucinations and visualize evaluation metrics for clear interpretation
• Tune hyperparameters and improve inference speed/accuracy through PEFT (LoRA/QLoRA) and advanced prompt engineering
• Collaborate with Product and Data Engineering teams to seamlessly integrate LLM features into the broader ecosystem
• Track and report progress using industry-standard benchmarks (MMLU, HumanEval, etc.) and custom internal KPIs
• Stay at the forefront of the field (e.g., State Space Models, new Transformer variants) and evaluate cutting-edge techniques for production readiness
• Engage in continuous technical growth and mentor junior colleagues to elevate the team's expertise
Qualifications and Skills
• 3+ years of commercial experience in Machine Learning, with a specific focus on the NLP or LLM domain
• Strong knowledge of Python3, NumPy, pandas, and modern text-processing libraries, PyTorch and Hugging Face (Transformers, PEFT, Accelerate)
• Proficiency in PEFT/LoRA and Reinforcement Learning techniques
• Deep understanding of attention mechanisms, tokenization, context window management, and embedding spaces
• Practical experience in at least one of the following: Retrieval-Augmented Generation (RAG), Fine-tuning, or Agentic frameworks
• Proven ability to manage and analyze massive datasets (>100GB) across text, image, and audio formats
• Hands-on experience crafting high-fidelity datasets and building robust data pipelines
• Expertise in prompt engineering, agentic framework design, and LLM pipeline orchestration
• Experience deploying LLMs to production environments using Triton Inference Server, vLLM, TGI, or ONNX
• Good written and spoken English
Nice to have
• Practical experience with Pinecone, Weaviate, Milvus, or Chroma
• Advanced quantization (GGUF, AWQ, EXL2), pruning, and knowledge distillation
• Experience with LangChain, LlamaIndex, or AutoGen
• Basic understanding of web/client-server architecture and streaming API responses (Asyncio, aiohttp)
• Familiarity with RAGAS, DeepEval, or G-Eval
• Experience using Docker, Kubernetes, and cloud GPU orchestration (e.g., Run:ai, Lambda Labs)
• Knowledge of C++, Triton, or CUDA for custom kernel development
We offer multiple benefits that include
• The environment of equal opportunities, transparent and value-based corporate culture, and an individual approach to each team member
• Competitive salary packages with performance-based annual reviews
• Employment via Contract of Employment (UoP) in complete alignment with Polish Labour Law
• Guaranteed paid vacation, public holidays, and medical leaves as per statutory regulations
• Continuous growth and development opportunities through internal knowledge hubs, corporate courses, and free English classes
• Comprehensive private medical insurance to supplement your standard NFZ coverage.
Our distributed team is looking for an experienced Applied Scientist with a strong background in Large Language models to develop high-performance Generative AI features across Cloud and Edge environments.
Job Summary
In this role you will drive the transition from research to production by optimizing local inference through model compression and quantization for private, real-time Edge performance, while also engineering scalable RAG architectures and multi-agent systems for Cloud deployment. Your daily responsibilities encompass the full research lifecycle, including formulating hypotheses, generating synthetic datasets, fine-tuning LLMs, and validating safety and alignment, ultimately culminating in technical reports.
Responsibilities and Duties
• Design and implement advanced methods in prompt orchestration, fine-tuning (SFT/RLHF/DPO), and autonomous agentic workflows
• Curate high-quality training data from large-scale text and multi-modal sources
• Identify patterns in model hallucinations and visualize evaluation metrics for clear interpretation
• Tune hyperparameters and improve inference speed/accuracy through PEFT (LoRA/QLoRA) and advanced prompt engineering
• Collaborate with Product and Data Engineering teams to seamlessly integrate LLM features into the broader ecosystem
• Track and report progress using industry-standard benchmarks (MMLU, HumanEval, etc.) and custom internal KPIs
• Stay at the forefront of the field (e.g., State Space Models, new Transformer variants) and evaluate cutting-edge techniques for production readiness
• Engage in continuous technical growth and mentor junior colleagues to elevate the team's expertise
Qualifications and Skills
• 3+ years of commercial experience in Machine Learning, with a specific focus on the NLP or LLM domain
• Strong knowledge of Python3, NumPy, pandas, and modern text-processing libraries, PyTorch and Hugging Face (Transformers, PEFT, Accelerate)
• Proficiency in PEFT/LoRA and Reinforcement Learning techniques
• Deep understanding of attention mechanisms, tokenization, context window management, and embedding spaces
• Practical experience in at least one of the following: Retrieval-Augmented Generation (RAG), Fine-tuning, or Agentic frameworks
• Proven ability to manage and analyze massive datasets (>100GB) across text, image, and audio formats
• Hands-on experience crafting high-fidelity datasets and building robust data pipelines
• Expertise in prompt engineering, agentic framework design, and LLM pipeline orchestration
• Experience deploying LLMs to production environments using Triton Inference Server, vLLM, TGI, or ONNX
• Good written and spoken English
Nice to have
• Practical experience with Pinecone, Weaviate, Milvus, or Chroma
• Advanced quantization (GGUF, AWQ, EXL2), pruning, and knowledge distillation
• Experience with LangChain, LlamaIndex, or AutoGen
• Basic understanding of web/client-server architecture and streaming API responses (Asyncio, aiohttp)
• Familiarity with RAGAS, DeepEval, or G-Eval
• Experience using Docker, Kubernetes, and cloud GPU orchestration (e.g., Run:ai, Lambda Labs)
• Knowledge of C++, Triton, or CUDA for custom kernel development
We offer multiple benefits that include
• The environment of equal opportunities, transparent and value-based corporate culture, and an individual approach to each team member
• Competitive salary packages with performance-based annual reviews
• Employment via Contract of Employment (UoP) in complete alignment with Polish Labour Law
• Guaranteed paid vacation, public holidays, and medical leaves as per statutory regulations
• Continuous growth and development opportunities through internal knowledge hubs, corporate courses, and free English classes
• Comprehensive private medical insurance to supplement your standard NFZ coverage.
🔍 Dekoder Ogłoszenia
🔴
drive the transition from research to production
Oznacza to, że będziesz odpowiedzialny za przenoszenie modeli z fazy eksperymentalnej do gotowych do użycia produktów, co może wiązać się z presją czasu i koniecznością radzenia sobie z problemami produkcyjnymi.
🔴
full research lifecycle
Choć brzmi to kompleksowo, może oznaczać, że będziesz musiał zajmować się wszystkimi etapami, od pomysłu po raport, bez wyraźnego podziału na role.
🔴
curate high-quality training data
Może to oznaczać ręczne filtrowanie i przygotowywanie danych, co jest czasochłonne i często mniej ekscytujące niż praca nad samymi modelami.
🔴
visualize evaluation metrics for clear interpretation
Choć tworzenie wizualizacji jest ważne, może to również sugerować, że obecne metryki nie są wystarczająco czytelne i będziesz musiał poświęcić czas na ich poprawę.
🟡
stay at the forefront of the field
Oznacza to ciągłe uczenie się i śledzenie nowości, co jest standardem w IT, ale może też sugerować brak jasno zdefiniowanych ścieżek rozwoju w ramach firmy.