The First Workshop·AACL-IJCNLP 2026
Autonomous Research & Recursive Self‑Improvement
Co-located with AACL-IJCNLP 2026 in Hengqin, China, on Tuesday, November 10, 2026.
Share your work on autonomous research and self-improving AI.
Submit by September 30, 2026, 11:59 PM AoE.
Abstract
AI agents are crossing a consequential threshold: from assisting with isolated research steps to running closed loops of literature search, hypothesis generation, coding, experimentation, evaluation, and revision. Whether those loops compound into recursive self-improvement, a claim often tied to artificial superintelligence, is prominent but unsettled, and it calls for measurement rather than assumption.
ARRSI asks what can actually be measured. We treat as open the claims most often made on an agent’s behalf: that its hypotheses are genuinely novel, that its gains reflect research ability rather than benchmark overfitting or undisclosed human labor, and that improvements derived from its own runs persist. Each becomes tractable only through reportable evidence: held-out results, attribution of a gain to the component that changed, budgets for compute and human intervention, and controls for contamination and reward tampering.
- Venue
- AACL-IJCNLP 2026
- Location
- Hengqin, China
- Date
- Tuesday, November 10, 2026
- Format
- One day, in person
- Submissions
- OpenReview · double-blind
About the Workshop
Why nowThe question is no longer speculative. Research agents now synthesize literature, propose hypotheses that have been taken into laboratory testing, and run closed research loops: proposing an idea, implementing it, running an experiment, evaluating the result, and choosing the next attempt. What this establishes is contested: it shows that generated hypotheses can be worth testing, not that novelty has been measured.
Self-improvement also reaches training: models generate practice data, rewards, curricula, and critiques through self-play, self-reward, and verifiable task generation, and agentic data systems improve the datasets they create. This broadens AI-for-AI to data curation, pretraining, midtraining, post-training, and systems optimization. Yet full recursive self-improvement remains unproven; evaluation must expose overfitting, hidden human labor, and reward tampering.
ARRSI fits AACL-IJCNLP because language is how research agents form hypotheses, retrieve multilingual evidence, call tools, critique claims, and explain results. The workshop brings together NLP, agentic systems, AI for Science, AI for AI, evaluation, and governance to establish definitions, benchmarks, reporting standards, and safety boundaries for systems that do not merely produce research-like text, but accumulate auditable evidence and use it to become better researchers.
Call for Papers
Empirical · theoretical · systems · dataset · benchmark · survey · positionWe invite empirical, theoretical, systems, dataset, benchmark, survey, and position papers on the following topics.
T1RSI for Harness & Loop Engineering
Prompts, context, memory, tools, workflows, evaluators, and harness code.
T2Autonomous Research
Literature synthesis, hypotheses, experiments, analysis, replication, peer review, and research taste.
T3AI for Science
Mathematics, physics, chemistry, biology, medicine, and materials science.
T4AI for AI
Models, agents, architectures, objectives, optimizers, kernels, compilers, training, inference, alignment, and evaluation.
T5Autonomous Data Engineering
Discovery, collection, cleaning, filtering, annotation, synthesis, provenance, and quality control.
T6Automated Training & Self-Training
Pretraining, midtraining, post-training, self-play, rewards, curricula, and continual learning.
T7RSI for Coding & General Agents
Coding and general agents, plus finance, healthcare, medicine, education, and other domains.
T8Infrastructure & Evaluation
Experiment managers, sandboxes, verifiers, provenance, checkpoints, rollback, compute accounting, benchmarks, and contamination control.
T9Safety & Governance
Validity, attribution, negative results, reward hacking, oversight, security, dual use, and human control.
| Important dates | |
|---|---|
| Call for papers released | September 1, 2026 |
| Submission deadline | September 30, 2026 |
| Notification of acceptance | October 7, 2026 |
| Camera-ready papers due | October 12, 2026 |
| Proceedings due (organizers) | October 15, 2026 |
| Workshop day | Tuesday, November 10, 2026 |
Schedule
Provisional · one day · breaks follow the conference| Time | Session |
|---|---|
| 09:00 | Opening and invited talks |
| 11:00 | Contributed talks, posters, and demonstrations |
| 14:00 | Further contributed and invited talks |
| 16:00 | Evaluation and reproducibility clinic |
| 17:00 | Synthesis: vocabulary, checklist, and research agenda |
Invited Speakers
To be announced.
We are inviting speakers across harness engineering, scientific agents, AI for Science, and the measurement and governance of recursive self-improvement. Confirmed speakers will appear here.
Organizers
General & program chairs, proceedings, website, D&I, logistics
Yuxuan Zhang↗
University of British Columbia · Vector Institute

Kelsey Allen↗
University of British Columbia · Vector Institute

Peter West↗
University of British Columbia

Yizhi Li↗
University of Manchester

Rui Meng↗
Google AI Research

Hanqi Yan↗
King’s College London

Hao Zhang↗
NVIDIA Research

Wenhu Chen↗
Meta Superintelligence Labs · University of Waterloo

Mingchen Zhuge↗
Recursive Superintelligence

Jian Yang↗
M-A-P

Xianglong Liu
M-A-P

Ming Zhou↗
Langboat

Dingmin Wang↗
Amazon AWS AI Lab
The organizing team spans North America, Europe, Asia, and the Middle East. Conflicted organizers take no part in assignments, discussion, or decisions.
Advisory Board

Yixuan He↗
Arizona State University

Robin Ding↗
UCLA

Xingwei Qu↗
University of Manchester

Terry Yue Zhuo↗
Monash University · CSIRO Data61

Ge Zhang↗
TokenWave.AI · M-A-P

Yilun Du↗
Assistant Professor · Harvard University

Zifan Wang↗
Meta · London

Baishakhi Ray↗
Associate Professor · Columbia University, USA

Luyu Gao↗
OpenAI · USA

Zhuowen Tu↗
Professor · University of California, San Diego

Yiran Chen↗
Professor · Duke University
Program Committee
Reviewing runs on OpenReview with double-blind submission, discussion, and decisions. We target at least 60 reviewers so that every submission receives three reviews while individual loads stay at three papers or fewer. The reviewer list will be announced once finalized. Reviewers assess evidence traceability, controls, reproducibility, resources, safety, and limitations; undisclosed AI-generated reviews are prohibited.
Submit Your Work
Submissions are now open on OpenReview. The submission deadline is September 30, 2026, 11:59 PM AoE.
Long papers run up to 8 pages and short, position, or resource papers up to 4 pages, excluding references, in *ACL format. Submissions are double-blind and will be reviewed on OpenReview. Every accepted paper presents a poster, and 4 to 6 receive talk slots. We welcome empirical, theoretical, systems, dataset, benchmark, survey, and position papers. Bring your findings, tools, benchmarks, and carefully documented negative results to the ARRSI community.
We plan to submit the final camera-ready papers to the ACL Anthology proceedings by October 15, 2026.