
What happened
The AI Security Lab outlined its standard experiment setup: Python programs call Ollama's API directly to run LLMs, and the code lives in the public GitHub repo kondonator/ai-security-playground.
Why it matters
Identical prompts can produce different answers depending on the model and other conditions, so the Lab treats each result as valid only under the specific conditions tested.
What to watch
The setup is a reference environment, and the Lab notes that Web UI use for running experiments is a future possibility. The next article will actually trigger Prompt Injection to observe the behavior.
WHO IT HITSSecurity researchers and developers who want to reproduce LLM security experiments, such as Prompt Injection tests, can now inspect and run the same public code and environment.
Summaries like this, in your inbox every morning.
This Zenn article is part of the AI Security Lab series. The previous article explained why the Lab was started and what it aims to clarify about how far work can be delegated to AI. This article focuses on the shared experiment environment that will be used in future articles, separating that explanation from the actual experiment results, which will be covered next.
The Lab uses Ollama to run LLMs locally, and experiments are executed by Python programs that call the Ollama API directly. The code is managed by experiment number under an experiments directory, and the whole setup is public in the GitHub repository kondonator/ai-security-playground. This is meant to make experiments reproducible and comparable by keeping conditions as consistent as possible.
The Lab stresses that results depend on many conditions, including the model, so a finding from one model does not automatically hold for all LLMs, and the same model may not give the same result under different conditions. Because of that, the Lab treats results as observations tied to specific conditions rather than absolute properties, and it is not using multiple models to rank them. The next article is set to move from environment setup to actual experiments, starting with Prompt Injection. Whether that demonstration will generalize beyond the Lab's environment is likely to depend on the specific models and conditions used, and the Lab's approach suggests readers should watch for those details in the upcoming results.
Pick your industry and the AI tools you use, and get news related to your work every day.
Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.
A 2026 Pew report found 10% of Americans use chatbots for emotional support or companionship, and Anthropic re…

The author built a palm-sized car where a PC runs a PyTorch CNN (NVIDIA PilotNet, shrunk) that maps camera ima…

Splice CEO Kakul Srivastava said she is "careful about AI-written documents" because when you cannot tell whet…

David Robinson, who wrote the safety reports accompanying every major OpenAI model release, resigned this week…

Instinct raised $1 billion in September 2026 at a $10 billion valuation, and a growing list of rivals — includ…

NVIDIA posted $89.02 billion in Data Center revenue for the second quarter of fiscal 2027, reported August 26…
