Gert LekGert Lek

> About

I am a second-year PhD researcher at the Université de Neuchâtel, supervised by Prof. Lydia Y. Chen. My research focuses on AI Safety & Alignment, computer vision, and NLP.

// recent news

  • [2026]Safety Reconstructed (LLaDA-Guard) accepted as a NeurIPS 2026 Spotlight - arXiv
  • [2026]OT Activation Steering accepted at ICLR 2026 TTU Workshop - read more
  • [2026]Detective SAM accepted at ICLR 2026 - read more
  • [2025]Detective SAM workshop paper accepted at ICML 2025 DIG-BUGS Workshop - read more

> Publications

Safety Reconstructed: Generative Modeling via Masked Diffusion Builds Strong Safety Guardrails

NeurIPS 2026 · Spotlight

Gert Lek, Abele Malan, Chaoyi Zhu, Pin-Yu Chen, Robert Birke, Lydia Y. Chen

Moves guard models from discriminative to generative classification to create a more robust training objective: rather than predicting a verdict token, LLaDA-Guard asks which safety label better explains the text, spreading supervision over every moderated token. It fine-tunes LLaDA-8B-Instruct with a class-conditional reconstruction objective, leading on average rank across seven held-out safety benchmarks with better calibration.

Camera-ready version and full code coming soon.

Paper

Understanding Confabulation and Rethinking Reconstruction in Activation Explanations

arXiv preprint 2026

Gert Lek, Zixuan Xia, Pin-Yu Chen, Lydia Y. Chen

Shows that as Natural Language Autoencoder explanations of model activations get better at predicting behavior, they increasingly contain unsupported details and writing flaws. Introduces an evaluation framework for information recovery, contextual support, and writing quality, and proposes Flow-NLA, which uses diffusion likelihood bounds to retain the utility gains of point reconstruction while curbing confabulation. Evaluated on Qwen, Gemma, and Apertus.

Paper

An Optimal Transport View of Activation Steering In Masked Diffusion Models

ICLR TTU Workshop 2026

Gert Lek, Chaoyi Zhu, Pin-Yu Chen, Robert Birke, Lydia Y. Chen

Introduces an Optimal Transport view of activation steering for masked diffusion models, learning a lightweight affine map that transports activation distributions between behaviors. Unifies common steering rules as special cases and improves instruction-following accuracy across LLaDA-Instruct, LLaDA 1.5, and Dream-Instruct.

Detective SAM: Adaptive AI-Image Forgery Localization

ICLR 2026

Gert Lek, Nicolas van Schaik, Chaoyi Zhu, Pin-Yu Chen, Robert Birke, Lydia Y. Chen

Extends SAM2 for image forgery localization with perturbation-driven forensic embeddings, lightweight feature adapters, and a learnable prompt module. Introduces AutoEditForge, an automated diffusion edit generation pipeline. Achieves SOTA on common benchmarks and modern editing models such as NanoBanana and Qwen-Image-Edit.

Detective SAM: Adapting SAM to Localize Diffusion-based Forgeries via Embedding Artifacts

ICML DIG-BUGS Workshop 2025

Gert Lek, Chaoyi Zhu, Pin-Yu Chen, Robert Birke, Lydia Y. Chen

Extends SAM with blur-driven forensic embedding signals, hierarchical learnable prompts, and lightweight adapters for automatic forgery mask generation. Outperforms prior methods on MagicBrush and CoCoGlide.

> Experience

PhD Researcher @ Université de Neuchâtel

2025 – Present

Research on AI Safety & Alignment, computer vision, and NLP. Supervised by Prof. Lydia Y. Chen.

R&D @ Ortec Finance

2023 – 2024

Reinforcement learning, bandit problems, and software development.

> Education

MSc Quantitative Finance @ Erasmus School of Economics

2022 – 2023

Thesis: "Interpretable High-dimensional Continuous RL" - Grade: 9.0/10

BSc Mathematics @ TU Delft

2019 – 2022

Thesis: "Robust Optimal Classification Trees" - Grade: 9.0/10

> Posts

> Contact

Feel free to reach out if you're interested in collaborating or have questions about my research.