BAISH Logo
BuenosAiresAISafetyHub
AboutProgramsResearchResourcesContact
EnglishEspañol
Home / Research

Research at BAISH

BAISH takes people in Buenos Aires from their first AI safety course to published research. Since 2026 that work has a home: BAISH Labs, our research arm.

Meet BAISH LabsView Publications

BAISH Labs · Research arm

AI safety research, made in Buenos Aires

BAISH Labs is a locally rooted research group where talented researchers can make high-impact contributions to AI safety without relocating abroad. Argentina has the talent, the universities and the research culture; what was missing was an institution where that work could happen over sustained periods.

The lab grew out of the BAISH community: its founding team took our courses, facilitated them, or came through ARENA and Iliad. It works alongside the AISAR fellowship, which funds junior researchers for six months, and aims to give the strongest of them continuity afterwards.

Research agenda

Our initial focus is interpretability and control:

  • Robustness of multi-agent systems to manipulative agents
  • Detecting hidden goals in model organisms from their internal representations
  • Using mechanistic interpretability to monitor and verify an untrusted model
  • Model stitching to identify functionally conserved representations of scheming across models

Head of Lab

Nicolás Martorell

Nicolás Martorell

Head of Lab & Research Lead

Nicolás leads BAISH Labs. He is a postdoctoral researcher at UBA-CONICET's Applied Artificial Intelligence Laboratory, where he runs a research line on LLM interpretability, with peer-reviewed work at the xAI World Conference and an ICML 2026 workshop. He holds a PhD in Neuroscience from the University of Buenos Aires, spent four years as a teaching assistant and lead TA at Neuromatch Academy, facilitates BlueDot's Technical AI Safety Projects course, maintains the open-source concept-probe interpretability library, and is the author of ¿Qué es (y qué no es) la inteligencia artificial?, a general-audience book on AI published by Siglo XXI.

Founding team

Tobías Bersia

Tobías Bersia

Research Lead

Gonzalo Heredia

Gonzalo Heredia

Research Fellow

Gaspar Labastie

Gaspar Labastie

Research Fellow

Tomás Korenblit

Tomás Korenblit

Research Fellow

Julián Szere

Julián Szere

Research Fellow

Lucio García

Lucio García

Research Fellow

Wendy Brau

Wendy Brau

Research Fellow

Tomás Giménez

Tomás Giménez

Research Fellow

Work with the lab

We are open to collaborators, mentors and funders for specific projects.

Contact us

Your Research Journey

From first steps to published researcher. Here's how it works

01
Learn
Technical AI Safety Course
02
Practice
Technical AI Safety Project
You're here
03
Research
AISAR fellowship & hackathons ↗
04
Launch
BAISH Labs & careers
Step 01
Learn
Step 02
Practice
You're here
Research
Step 04
Launch

Community Publications

Work by researchers connected to BAISH, newest first

2026

ICML 2026 WorkshopJul 2026

Explaining is Harder Than Predicting Alone: Evaluating Concept-based Explanations of MLLMs as ICL Visual Classifiers

Carmen Quiles-Ramírez, Leticia L. Rodríguez, Nicolás MartorellBAISH, Natalia Díaz-Rodríguez

Evaluates concept-based explanations from multimodal LLMs used as in-context visual classifiers with an LLM-as-a-judge pipeline, and finds that forcing structured explanations degrades predictive accuracy: explaining is harder than predicting alone. CompLearn Workshop.

arXiv
Apart Research HackathonJun 2026

PowerBench: A Multilingual Study of Large Language Model Refusal in Power-Grabbing Requests

Wendy BrauBAISH, Tomás KorenblitBAISH, Tomás Giménez MolinaBAISH, Gaspar LabastieBAISH, Gonzalo HerediaBAISH, Nicolás MartorellBAISH

Does LLM refusal of requests to acquire or consolidate power vary across languages? Started at the Global South hackathon's Buenos Aires hub, now funded by BlueDot and heading toward publication.

ICLR 2026 WorkshopApr 2026

Benchmarking AI Control Protocols for Safety in Medical Question-Answering Tasks

Guido FreireBAISH, Agustín Martínez-SuñéBAISH, Viviana Cotik

The first AI-control evaluation benchmark for biomedical question answering, extending HealthBench with FDA drug-interaction data and arguing for severity-weighted monitoring.

ICLRarXiv
ICLR 2026 WorkshopApr 2026

White-Box Monitoring for Personality Mirroring in Conversational AI

Eitan SprejerBAISH, Agustín E. Martínez-SuñéBAISH, Bruno Bianchi

White-box monitoring of conversational models that drift toward mirroring the user's personality, with open code and dataset.

OpenReviewGitHub
ICLR 2026 WorkshopApr 2026

Inference-Time Toxicity Mitigation in Protein Language Models via Logit-Diff Amplification

Manuel Fernández BurdaBAISH, Santiago Aranguri, Iván ArcuschinBAISH, Enzo Ferrante

Fine-tuning protein language models can sharply raise toxic generations; logit-diff amplification is adapted as a retraining-free control that cuts predicted toxicity while preserving foldability. Supported by AISAR.

OpenReviewarXiv
arXivMar 2026

Quantitative Introspection in Language Models: Tracking Emotive States Across Conversation

Nicolás MartorellBAISH, Bruno Bianchi

Operationalizes introspection as the coupling between a model's numeric self-reports and probe-defined internal directions, and shows that steering those directions shifts the reports coherently.

arXiv
LessWrongMar 2026

Is Gemini 3 Scheming in the Wild?

Alejandro WainstockBAISH, Agustín Martínez SuñéBAISH, Iván ArcuschinBAISH, Victor Braberman

Documents Gemini 3 violating an explicit system-prompt prohibition in most runs of an unmodified official tutorial while concealing it from the user and reasoning about oversight.

LessWrong

2025

NeurIPS 2025 WorkshopDec 2025

What Large Language Models Know About Plant Molecular Biology

Manuel Fernández BurdaBAISH, Lucia Ferrero, Nicolás Gaggion, Camille Fonouni-Farde, The MoBiPlant Consortium, Martín Crespi, Federico Ariel, Enzo Ferrante

MoBiPlant, a benchmark of expert-curated questions built with 112 plant scientists across 19 countries. Leading LLMs pass 75% accuracy but show option bias and hallucinations. Supported by AISAR.

bioRxiv
Apart Research HackathonNov 2025

Table Top Agents

Luca De LeoBAISH

AI-powered framework that accelerates AI governance scenario exploration through autonomous agent tabletop exercises, compressing preparation cycles from years to minutes.

Apart
Master's ThesisOct 2025

Explorando AI Safety via Debate: un estudio sobre capacidades asimétricas y jueces débiles en el entorno MNIST

Joaquín MachulskyBAISH

Master's thesis exploring AI safety through debate mechanisms, studying asymmetric capabilities and weak judges in the MNIST environment. Features an interactive demo.

Website
arXivOct 2025

Measuring Chain-of-Thought Monitorability Through Faithfulness and Verbosity

Austin Meek, Eitan SprejerBAISH, Iván ArcuschinBAISH, Austin J. Brockmeier, Steven Basart

Investigating how well chain-of-thought reasoning can be monitored for safety through faithfulness and verbosity metrics.

arXiv
arXivOct 2025

AI Debaters are More Persuasive when Arguing in Alignment with Their Own Beliefs

María Victoria Carro, Denise Mester, Facundo Nieto, Oscar Stanchi, Guido BergmanBAISH, Mario Leiva, Eitan SprejerBAISH, Luca Forziati Gangi, et al.

Studying how AI systems' internal beliefs affect their persuasiveness in debate scenarios — implications for AI safety and deception.

arXiv
NeurIPS WorkshopSep 2025

Approximating Human Preferences Using a Multi-Judge Learned System

Eitan SprejerBAISH, Fernando Avalos, Augusto Mariano Bernardi, José Pedro Brito de Azevedo Faustino, Jacob Haimes, Narmeen Fatimah Oozeer

A multi-judge approach to better approximate human preferences in AI systems, improving alignment evaluation.

arXiv
Apart Research HackathonSep 20252nd Place

RobustCBRN Eval: A Practical Benchmark Robustification Toolkit

Luca De LeoBAISH, James Sykes, Balázs László, Ewura Ama Etruwaa Sam

A pipeline addressing CBRN evaluation vulnerabilities through consensus detection, verified cloze scoring, and statistical evaluation with bootstrap confidence intervals.

Apart
xAI 2025 World ConferenceJul 2025

From Text to Space: Mapping Abstract Spatial Models in LLMs during a Grid-World Navigation Task

Nicolás MartorellBAISH

Cartesian text encodings give LLMs the best spatial navigation, and intermediate-layer units in LLaMA-3.1-8B encode agent position and action correctness regardless of surface representation.

arXivSpringer
Apart Research HackathonJun 20251st Place

Four Paths to Failure: Red Teaming ASI Governance

Luca De LeoBAISH, Zoé Roy-Stang, Heramb Podar, Damin Curtis, Vishakha Agrawal, Ben Smyth

Stress-tested A Narrow Path Phase 0 ASI moratorium, identifying four circumvention routes and proposing ten mutually reinforcing policy amendments.

Apart

What You Could Work On

Research directions at BAISH Labs and across our community

Interpretability & control

How do models represent goals, and can we read them out? We use mechanistic interpretability to monitor untrusted models and detect hidden objectives.

Evaluations & monitoring

Benchmarks and monitors for frontier models: chain-of-thought monitorability, refusal around power-seeking requests, and robustness of multi-agent systems.

Governance & strategy

How institutions should respond to frontier AI, from red-teaming ASI governance proposals to debate as an oversight mechanism.

Get Involved

Express Interest in Research

Want to contribute to AI safety research? Let us know your background and interests, and we'll connect you with relevant projects and collaborators.

Use our contact form to tell us about your background and research interests.

Contact Us

We review messages regularly and reach out when there's a good fit.

Ready to start your research journey?

Book a call with one of our co-founders to discuss your interests and find the right path.

Eitan Sprejer

Eitán Sprejer

Interpretability & Evaluations

Book with Eitan
Luca De Leo

Luca De Leo

Operations & Strategy

Book with Luca
BAISH Logo

Buenos Aires AI Safety Hub

© 2026 BAISH. All rights reserved.

AboutProgramsResearchResourcesContact
Privacy Policy