AIToday
Large Language ModelsAI Coding AssistantsOpen-Source AIZenn AI/MLPublished: Oct 3, 2026, 22:00 JST

AI Security Lab opens Ollama-based LLM test setup on GitHub

AI Security Lab opens Ollama-based LLM test setup on GitHub

3 Key Points

  1. What happened

    The AI Security Lab outlined its standard experiment setup: Python programs call Ollama's API directly to run LLMs, and the code lives in the public GitHub repo kondonator/ai-security-playground.

  2. Why it matters

    Identical prompts can produce different answers depending on the model and other conditions, so the Lab treats each result as valid only under the specific conditions tested.

  3. What to watch

    The setup is a reference environment, and the Lab notes that Web UI use for running experiments is a future possibility. The next article will actually trigger Prompt Injection to observe the behavior.

WHO IT HITSSecurity researchers and developers who want to reproduce LLM security experiments, such as Prompt Injection tests, can now inspect and run the same public code and environment.

Not sure about something? Ask the AI

Questions and answers are published on this page.

Summaries like this, in your inbox every morning.

Context & Analysis

This Zenn article is part of the AI Security Lab series. The previous article explained why the Lab was started and what it aims to clarify about how far work can be delegated to AI. This article focuses on the shared experiment environment that will be used in future articles, separating that explanation from the actual experiment results, which will be covered next.

The Lab uses Ollama to run LLMs locally, and experiments are executed by Python programs that call the Ollama API directly. The code is managed by experiment number under an experiments directory, and the whole setup is public in the GitHub repository kondonator/ai-security-playground. This is meant to make experiments reproducible and comparable by keeping conditions as consistent as possible.

The Lab stresses that results depend on many conditions, including the model, so a finding from one model does not automatically hold for all LLMs, and the same model may not give the same result under different conditions. Because of that, the Lab treats results as observations tied to specific conditions rather than absolute properties, and it is not using multiple models to rank them. The next article is set to move from environment setup to actual experiments, starting with Prompt Injection. Whether that demonstration will generalize beyond the Lab's environment is likely to depend on the specific models and conditions used, and the Lab's approach suggests readers should watch for those details in the upcoming results.

FAQ
Who can use this experiment environment?
Researchers and developers who want to reproduce LLM security experiments, such as Prompt Injection tests, can inspect and run the same public code from the GitHub repository.
Which LLM models are used in the experiments?
The Lab uses multiple models, including gpt-oss:20b and gemma4:12b, to compare results under the same conditions.
Will the same prompt always produce the same result?
No, the Lab says results can differ because of the model, prompt, system prompt, input data, conversation history, inference settings, and execution environment.

AI news that matters for your work, delivered every morning.

Pick your industry and the AI tools you use, and get news related to your work every day.

Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.

Questions and answers are published on this page.

Related Articles

Next articleKaran Joshi extracts Muse files on everyone you know