Data Scientist

Itamar Efrati

I build machine learning systems end to end, from raw data and preprocessing through modeling and optimization to pipelines that run in production. My work spans computer vision, time series and signals, video, and agentic systems.

Published at ICMLA 2024, HICSS 2024 and CHIL 2025 · M.Sc. Computer Science, Reichman University

About

My name is Itamar. I'm a Data Scientist with a background in research and software engineering. In recent years I've focused on building Machine Learning and Deep Learning solutions end to end: from raw data and preprocessing, through feature engineering, modeling, and hyperparameter optimization, to scalable pipelines and production systems. I've worked with both classical machine learning and deep learning across a wide range of data types, including structured data, time series, signals, images, and video, primarily using Python, PyTorch, PyTorch Lightning, and scikit-learn.

Before moving into Data Science, I spent several years as a software developer at the Prime Minister's Office, building internal systems, ETL pipelines, and infrastructure for a large data cluster, working with Python, C#, SQL, and a range of monitoring tools.

Alongside that, I built up substantial research experience. During my M.Sc. I led research on analyzing cellphone signals to predict chronic pain and suicide risk. The work involved reading and implementing academic papers, developing new models, and scientific writing, and led to publications at ICMLA, HICSS, and CHIL.

Today I work as a Data Scientist at Teva Pharmaceuticals. I've worked on a range of projects in healthcare and drug development. The main ones were computer vision systems for analyzing biological images, real-time video monitoring systems for assessing patient severity, and LLM-based and agentic solutions for automating processes and generating reports.

I'm looking to keep growing in a strong technological environment, working on complex problems with real impact, and to go deeper into AI, Machine Learning, and production systems. It matters to me to build solutions that don't stop at the POC or research stage, but ones that can actually be used, effectively and efficiently, with sensible use of resources and a proper fit for production environments.

Projects

Eight projects across industry, research, and personal AI engineering.

BioScreen: automated cell-line scoring

In production

Built and deployed a computer vision system that scores cell-line candidates from microscopy images, helping scientists select high producers for drug development.

Computer Vision Image Segmentation OpenCV Optuna Production Pipeline

Production deployment · ROC-AUC 0.734 Automated scoring to support cell-line selection

Tardive Dyskinesia severity from patient video

Clinical POC

Developed a clinical proof of concept that assesses involuntary movement severity from patient videos, translating visual movement into a quantitative score aligned with clinician assessments.

Video MediaPipe Signal Processing Time Series Clinical Biomarker

ρ = 0.80 agreement with the in-site clinician A human remote rater watching the same video reaches 0.57

Clinical Data Science Agent

Agentic

Built an agent that carries clinical trial analysis from raw data to modeling and statistical conclusions, accelerating the workflow while keeping the data scientist in control of key decisions.

AI Agents Skill Design Prompt Engineering Statistical Inference Human-in-the-Loop

Turns about a month of analysis into several days A coworker with clinical knowledge that helps design experiments and validate analyses.

ReportGenerator: automated validation reports

Agentic

Automated the creation of pharmaceutical validation reports from laboratory data and regulatory protocols, bringing calculations, results, and narrative into a complete report.

AI Agents ETL Document Generation Python OOP

Days or weeks per report, down to minutes Orchestrates report creation, from extracting data to composing the final report.

Research · M.Sc., Reichman University

Two-step digital phenotyping

HICSS 2024 · CHIL 2025

Developed methods for predicting suicide risk and chronic-pain health states from passive cellphone data despite scarce clinical labels, contributing the methodology to two published studies.

Clustering Feature Engineering Digital Health scikit-learn

Predicting health risks from cellphone usage Introduced a new classical machine learning method for time-series classification, applied to suicide-risk and chronic-pain prediction with limited clinical labels.

GP-VIB: information bottleneck with GP priors

ICMLA 2024

Developed a probabilistic time-series classifier that outperformed GP-VAE on Healing MNIST and achieved the best average rank among deep models in the UEA evaluation.

PyTorch Variational Inference Gaussian Processes Time Series

State-of-the-art time-series classification model Achieved the best average rank among evaluated deep models on UEA and 88.2% accuracy on Healing MNIST, compared with 77.9% for GP-VAE.

SS-GP-VIB: semi-supervised extension

M.Sc. thesis

Extended GP-VIB to learn from both labeled and unlabeled time series, improving classification when labels are scarce and observations are missing.

Semi-Supervised PyTorch Lightning Representation Learning

Outperforms GP-VIB when labeled data is limited Uses both labeled and unlabeled time series to improve classification accuracy, with larger gains when labels are scarce and observations are missing.

Personal · AI Engineering

Personal AI System

Personal project

Built a private family coordination system that brings messages, tasks, files, and calendar events into shared plans and daily routines.

AI Agent Engineering Prompt Architecture Tool Use Safe Intake Flask

A household system that automatically extracts tasks from incoming information Turns incoming messages and documents into tasks, shared plans, and daily routines.

Experience

2021–Present
Data Scientist · Teva Pharmaceuticals
  • Computer vision on large unlabeled biological image sets for drug development
  • Real-time video-based patient severity monitoring
  • LLM-based and agentic systems for analysis automation and report generation
  • End-to-end feature extraction and training pipelines on GPU clusters
2018–2021
Software Developer · Prime Minister's Office
  • ETL pipelines and infrastructure for a large data cluster
  • Python, C# and SQL across internal services
  • Monitoring and alerting with Grafana, Zabbix, Prometheus and Elasticsearch

Education

2025
M.Sc. Computer Science · Reichman University
GPA 93.9 · supervised by Dr. Shai Fine
2020
B.Sc. Computer Science · Reichman University
GPA 92.5

Publications

HICSS 2024
Predicting Adolescent Suicide Risk From Cellphone Usage Data and Self-Report Assessments
Responsible for the methodology
ICMLA 2024
Variational Information Bottleneck with Gaussian Process Priors for Time Series Classification
Efrati & Fine · First author
CHIL 2025
Predicting Health States of Patients with Chronic Pain from Cellphone Usage Data
PMLR 287 · Responsible for the methodology
M.Sc. 2025
Variational Information Bottleneck with Gaussian Process Priors for Semi-Supervised Time Series Classification
M.Sc. thesis · Reichman University

Skills

Machine Learning
Deep Learning Classical ML Semi-Supervised Learning Variational Inference Gaussian Processes Probabilistic Modeling Feature Engineering Hyperparameter Optimization Imbalanced & Scarce Labels Model Evaluation Explainability
Computer Vision
Image Segmentation Contour & Morphology Analysis Facial Landmark Tracking Video Processing Biological Imaging Intensity & Texture Features
Time Series & Signals
Time Series Classification Signal Processing Frequency & FFT Features Windowing & Segmentation Missing & Irregular Data Sensor & Usage Data
Agentic AI & LLMs
AI Agent Engineering Skill-Based Orchestration Prompt Engineering Human-in-the-Loop Design Agent Memory Design Tool Use LLM Text Generation Guardrails & Error Policy Config-Driven Agents
Python
PyTorch PyTorch Lightning scikit-learn OpenCV scikit-image MediaPipe SciPy NumPy Pandas Optuna XGBoost statsmodels SHAP aeon Matplotlib / Seaborn
Engineering
Production Pipelines Producer / Consumer Architectures Multithreading Hydra / OmegaConf pytest Model Registry GPU Clusters ETL SQL C# Git Grafana Elasticsearch
Research
Scientific Writing Paper Implementation Experimental Design Peer-Reviewed Publication
Languages
Hebrew · native English · professional