
Researchers at IIT Bombay and Adobe Research have developed a technique that reconstructs the original prompts users fed into large language models by analyzing only the output text, achieving near-perfect accuracy without needing access to the model itself.
The method works even across different models—an inverse model trained on a small open chatbot can extract prompts from GPT-4o responses—creating a security vulnerability that could expose proprietary system instructions and user queries.
What happened
Researchers at IIT Bombay and Adobe Research developed a method called "Previous-Token Prediction" (PTP) that reconstructs prompts fed to large language models by training an inverse model on synthetically generated data. The approach works without access to model weights and can reconstruct exact prompts or generate multiple semantically similar variants from a single LLM response.
Why it matters
The method exposes a broad security risk—companies risk exposing proprietary system prompts containing trade secrets and moderation rules, while individuals face extraction of personal or sensitive queries from text output. A small, open inversion model trained on one LLM (such as Qwen-3-0.6B) can reconstruct prompts from other models' responses (like GPT-4o), meaning attackers need not even know which model generated the text.
What to watch
The paper does not make explicit claims about attacks on commercial systems, but if the method works on current production models, AI labs will need to address and patch the issue quickly.
Ask the AI about this article →
The research reveals a fundamental asymmetry in how language models work: while forward prediction (next token) is the standard method for generation, the inverse process (reconstructing prior tokens) was long assumed to be impractical because many different prompts can produce similar responses. The IIT Bombay and Adobe team have overturned that assumption by training a specialized inverse model on synthetic data, demonstrating that near-perfect reconstruction is feasible without any access to the original model's weights or internals.
The security implications are severe because the attack surface is broad and the barrier to entry is low. A company or individual cannot rely on keeping their model proprietary to protect prompt secrecy—the method works across model boundaries, as shown by the Qwen-trained inverse successfully extracting meaning from GPT-4o output. System prompts often encode critical business logic, content moderation rules, and specialized instructions that companies treat as trade secrets. Individual users face parallel risk: sensitive queries, personal information, or confidential research embedded in a prompt could leak through an innocuous-looking output. The paper itself refrains from claiming the method has been deployed against live commercial systems, but the authors' framing suggests that if it succeeds on current production models, the AI research community and industry will need to develop mitigations rapidly.
For example, today's edition would include:
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · takes 30 seconds · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. Q&As are published on this page for other readers too.
Visko raised $10 million in pre-seed funding from Llama Ventures and opened public access to its first foundat…
AI company Runway has unveiled Solaris, the first model in a new category it calls "Interface World Models." I…

Google's AI search gave advice to call emergency services for users alone with an African, Indian, or Pakistan…

John Deere introduced JD, a conversational AI tool that lets farmers ask open-ended questions about their hist…

Nvidia CEO Jensen Huang said on Fox Business that AI is creating 'hundreds of thousands' of jobs, including in…

Israeli startup DataAgent Ltd