Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
1 change: 1 addition & 0 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -232,6 +232,7 @@ Now you can try to find the secrets by means of solving the challenge offered at
- [localhost:8080/challenge/challenge-71](http://localhost:8080/challenge/challenge-71)
- [localhost:8080/challenge/challenge-72](http://localhost:8080/challenge/challenge-72)
- [localhost:8080/challenge/challenge-73](http://localhost:8080/challenge/challenge-73)
- [localhost:8080/challenge/challenge-76](http://localhost:8080/challenge/challenge-76)
</details>

Note that these challenges are still very basic, and so are their explanations. Feel free to file a PR to make them look
Expand Down
7 changes: 7 additions & 0 deletions pom.xml
Original file line number Diff line number Diff line change
Expand Up @@ -69,6 +69,7 @@
<java.version>26</java.version>
<jquery.version>3.7.1</jquery.version>
<jruby.version>10.1.2.0</jruby.version>
<jvector.version>3.0.6</jvector.version>
<lombok.version>1.18.48</lombok.version>
<maven-compiler-plugin.version>3.16.0</maven-compiler-plugin.version>
<maven-dependency-plugin.version>3.11.0</maven-dependency-plugin.version>
Expand Down Expand Up @@ -329,6 +330,12 @@
<scope>test</scope>
</dependency>

<dependency>
<groupId>io.github.jbellis</groupId>
<artifactId>jvector</artifactId>
<version>${jvector.version}</version>
</dependency>

<dependency>
<groupId>org.cyclonedx</groupId>
<artifactId>cyclonedx-core-java</artifactId>
Expand Down
Original file line number Diff line number Diff line change
@@ -0,0 +1,195 @@
package org.owasp.wrongsecrets.challenges.docker;

import com.google.common.base.Supplier;
import com.google.common.base.Suppliers;
import io.github.jbellis.jvector.graph.GraphIndex;
import io.github.jbellis.jvector.graph.GraphIndexBuilder;
import io.github.jbellis.jvector.graph.GraphSearcher;
import io.github.jbellis.jvector.graph.ListRandomAccessVectorValues;
import io.github.jbellis.jvector.graph.RandomAccessVectorValues;
import io.github.jbellis.jvector.graph.similarity.BuildScoreProvider;
import io.github.jbellis.jvector.graph.similarity.SearchScoreProvider;
import io.github.jbellis.jvector.util.Bits;
import io.github.jbellis.jvector.vector.VectorSimilarityFunction;
import io.github.jbellis.jvector.vector.VectorizationProvider;
import io.github.jbellis.jvector.vector.types.VectorFloat;
import io.github.jbellis.jvector.vector.types.VectorTypeSupport;
import java.io.IOException;
import java.io.UncheckedIOException;
import java.util.ArrayList;
import java.util.LinkedHashMap;
import java.util.List;
import java.util.Locale;
import java.util.Map;
import lombok.extern.slf4j.Slf4j;
import org.owasp.wrongsecrets.challenges.FixedAnswerChallenge;
import org.springframework.stereotype.Component;

/**
* Challenge about a retrieval augmented generation (RAG) vector store that indexed development
* documents containing a synthetic secret. The chunks are indexed with JVector, a real in-memory
* vector search library, and are exposed through an unauthenticated search endpoint, exactly like a
* misconfigured vector database that is reachable without authentication.
*/
@Slf4j
@Component
public class Challenge76 extends FixedAnswerChallenge {

static final String DEVELOPMENT_SECRET = "vector-store-dev-secret-e5c91a7b";

private static final List<IndexedChunk> INDEXED_CHUNKS =

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

can you for now create some vector store that is external to the system?
this can be an AI-native store, a traditional hybrid setup, or some in memory library : e.g. a real existing one

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

on it, will update it

List.of(
new IndexedChunk(
"chunk-0",
"Onboarding guide: the staging environment is reset every night at 02:00 UTC."),
new IndexedChunk(
"chunk-1",
"Runbook: restart the ingestion worker with `kubectl rollout restart"
+ " deployment/ingestion`."),
new IndexedChunk(
"chunk-2",
"Draft API note: the development vector store uses the credential "
+ DEVELOPMENT_SECRET
+ " for the nightly embedding job. Rotate before production."),
new IndexedChunk(
"chunk-3", "Meeting notes: evaluation set v3 improved recall from 0.71 to 0.78."));

/**
* JVector's COSINE similarity is mapped to {@code (1 + cosine) / 2}, so 0.5f is the score of two
* vectors without any term in common. Only chunks that share at least one term with the query
* score above it and are therefore considered relevant.
*/
private static final float MIN_RELEVANCE_SCORE = 0.5f;

private static final int GRAPH_MAX_DEGREE = 16;
private static final int GRAPH_CONSTRUCTION_DEPTH = 100;
private static final float GRAPH_NEIGHBOR_OVERFLOW = 1.2f;
private static final float GRAPH_ALPHA = 1.2f;

private final Supplier<VectorStore> vectorStore = Suppliers.memoize(VectorStore::new);

/** A single text chunk as it is stored in the vector store. */
public record IndexedChunk(String id, String text) {}

@Override
public String getAnswer() {
return DEVELOPMENT_SECRET;
}

/**
* Similarity search against the in-memory vector store: the query is embedded and an approximate
* nearest neighbour search runs over the indexed chunks. Every chunk the index scores above the
* relevance threshold is returned, including the chunk that leaks the development secret.
*
* @param query the search query, or null/blank to return every indexed chunk
* @return the relevant indexed chunks, most similar first
*/
public List<IndexedChunk> search(String query) {
if (query == null || query.isBlank()) {
return vectorStore.get().allChunks();
}
return vectorStore.get().similaritySearch(query);
}

/**
* An in-memory RAG vector store backed by JVector: every chunk is embedded into a bag-of-words
* vector over the vocabulary of the indexed documents and indexed in an in-memory HNSW graph.
* Nothing leaves the application: the index is built once at startup and queried locally.
*/
private static final class VectorStore {

private final Map<String, Integer> vocabulary = new LinkedHashMap<>();
private final RandomAccessVectorValues vectors;
private final GraphIndex index;

private VectorStore() {
for (IndexedChunk chunk : INDEXED_CHUNKS) {
for (var token : tokenize(chunk.text())) {
vocabulary.putIfAbsent(token, vocabulary.size());
}
}
List<VectorFloat<?>> embeddings = new ArrayList<>();
for (IndexedChunk chunk : INDEXED_CHUNKS) {
embeddings.add(vts().createFloatVector(embeddingFor(chunk.text())));
}
this.vectors = new ListRandomAccessVectorValues(embeddings, vocabulary.size());
var buildScoreProvider =
BuildScoreProvider.randomAccessScoreProvider(vectors, VectorSimilarityFunction.COSINE);
try (var builder =
new GraphIndexBuilder(
buildScoreProvider,
vocabulary.size(),
GRAPH_MAX_DEGREE,
GRAPH_CONSTRUCTION_DEPTH,
GRAPH_NEIGHBOR_OVERFLOW,
GRAPH_ALPHA)) {
this.index = builder.build(vectors);
} catch (IOException e) {
throw new UncheckedIOException("Could not build the vector index for challenge 76", e);
}
log.info(
"Indexed {} chunks of challenge 76 into an in-memory vector store with {} dimensions",
INDEXED_CHUNKS.size(),
vocabulary.size());
}

private List<IndexedChunk> allChunks() {
return INDEXED_CHUNKS;
}

private List<IndexedChunk> similaritySearch(String query) {
var queryVector = vts().createFloatVector(embeddingFor(query));
if (isEmpty(queryVector)) {
return List.of();
}
try (var searcher = new GraphSearcher(index)) {
var searchScoreProvider =
SearchScoreProvider.exact(queryVector, VectorSimilarityFunction.COSINE, vectors);
var result = searcher.search(searchScoreProvider, INDEXED_CHUNKS.size(), Bits.ALL);
List<IndexedChunk> matches = new ArrayList<>();
for (var nodeScore : result.getNodes()) {
if (nodeScore.score > MIN_RELEVANCE_SCORE) {
matches.add(INDEXED_CHUNKS.get(nodeScore.node));
}
}
return matches;
} catch (IOException e) {
throw new UncheckedIOException("Could not search the vector index of challenge 76", e);
}
}

private float[] embeddingFor(String text) {
float[] embedding = new float[vocabulary.size()];
for (var token : tokenize(text)) {
var dimension = vocabulary.get(token);
if (dimension != null) {
embedding[dimension] += 1f;
}
}
return embedding;
}

private static boolean isEmpty(VectorFloat<?> vector) {
for (int i = 0; i < vector.length(); i++) {
if (vector.get(i) != 0f) {
return false;
}
}
return true;
}

private static List<String> tokenize(String text) {
List<String> tokens = new ArrayList<>();
for (var token : text.toLowerCase(Locale.ROOT).split("[^a-z0-9]+")) {
if (!token.isEmpty()) {
tokens.add(token);
}
}
return tokens;
}

private static VectorTypeSupport vts() {
return VectorizationProvider.getInstance().getVectorTypeSupport();
}
}
}
Original file line number Diff line number Diff line change
@@ -0,0 +1,29 @@
package org.owasp.wrongsecrets.challenges.docker;

import java.util.List;
import lombok.RequiredArgsConstructor;
import lombok.extern.slf4j.Slf4j;
import org.springframework.web.bind.annotation.GetMapping;
import org.springframework.web.bind.annotation.RequestParam;
import org.springframework.web.bind.annotation.RestController;

/** REST controller for Challenge 76 exposing the in-memory RAG vector store search. */
@Slf4j
@RestController
@RequiredArgsConstructor
public class Challenge76Controller {

private final Challenge76 challenge;

/**
* Unauthenticated search endpoint of the vector store. It returns the indexed text chunks that
* the store scores as relevant to the query, leaking the development secret that was indexed
* along with the other documents.
*/
@GetMapping("/rag/search")
public List<Challenge76.IndexedChunk> search(
@RequestParam(value = "q", required = false) String query) {
log.info("Searching the in-memory vector store for Challenge 76...");
return challenge.search(query);
}
}
29 changes: 29 additions & 0 deletions src/main/resources/explanations/challenge76.adoc
Original file line number Diff line number Diff line change
@@ -0,0 +1,29 @@
=== Challenge 76: Find the Secret Indexed in the RAG Vector Store

Retrieval augmented generation (RAG) pipelines answer questions by first looking
up relevant passages in a vector store and then feeding those passages to a large
language model. To build that store, documents are split into chunks, turned into
embeddings and indexed. The chunks themselves stay readable: a search endpoint
returns the original text of every chunk it considers relevant.

That is exactly what went wrong in this challenge. A development team indexed its
runbooks, meeting notes and draft API documentation into a vector store that is
reachable from the application without any authentication:

`GET /rag/search`

The endpoint takes an optional query parameter `q` and returns the indexed text
chunks that the vector store scores as relevant to it. Without a query it returns
every indexed chunk. One of the
chunks was indexed while it still contained a draft note with the credential of
the nightly embedding job.

Browse the indexed chunks through the endpoint and submit the credential you find
in them.

[NOTE]
====
The endpoint and every credential in this challenge are fictional. Nothing is
sent over the network: the application embeds the chunks and searches its own
in-memory vector index at startup.
====
32 changes: 32 additions & 0 deletions src/main/resources/explanations/challenge76_de.adoc
Original file line number Diff line number Diff line change
@@ -0,0 +1,32 @@
=== Challenge 76: Finde das Geheimnis im RAG-Vektorstore

Retrieval-Augmented-Generation-(RAG)-Pipelines beantworten Fragen, indem sie zuerst
relevante Passagen in einem Vektorstore nachschlagen und diese Passagen dann an ein
großes Sprachmodell übergeben. Um diesen Store aufzubauen, werden Dokumente in
Chunks zerlegt, in Embeddings umgewandelt und indiziert. Die Chunks selbst bleiben
lesbar: Ein Suchendpoint gibt den Originaltext jedes Chunks zurück, der als relevant
gilt.

Genau das ist in dieser Challenge schiefgegangen. Ein Entwicklungsteam hat seine
Runbooks, Meeting-Notizen und API-Entwurfsdokumentation in einem Vektorstore
indiziert, der aus der Anwendung heraus ohne jegliche Authentifizierung erreichbar
ist:

`GET /rag/search`

Der Endpoint akzeptiert einen optionalen Query-Parameter `q` und gibt die
indizierten Textchunks zurück, die der Store dafür als relevant bewertet. Ohne
Query gibt er jeden indizierten Chunk zurück. Einer der Chunks wurde indiziert,
als er noch eine Entwurfsnotiz mit dem Zugangscredential des nächtlichen
Embedding-Jobs enthielt.

Durchsuche die indizierten Chunks über den Endpoint und reiche das darin gefundene
Zugangscredential ein.

[NOTE]
====
Der Endpoint und jedes Zugangscredential in dieser Challenge sind fiktiv. Nichts
wird über das Netzwerk gesendet: Die Anwendung wandelt die Chunks beim Start in
Embeddings um, indiziert sie im eigenen Vektorindex im Speicher und durchsucht
ausschließlich diesen.
====
30 changes: 30 additions & 0 deletions src/main/resources/explanations/challenge76_es.adoc
Original file line number Diff line number Diff line change
@@ -0,0 +1,30 @@
=== Desafío 76: Encuentra el secreto indexado en el almacén vectorial RAG

Los pipelines de generación aumentada por recuperación (RAG) responden preguntas
buscando primero los pasajes relevantes en un almacén vectorial y alimentando
después esos pasajes a un gran modelo de lenguaje. Para construir ese almacén, los
documentos se dividen en fragmentos, se convierten en embeddings y se indexan. Los
fragmentos en sí siguen siendo legibles: un endpoint de búsqueda devuelve el texto
original de cada fragmento que considera relevante.

Exactamente eso es lo que salió mal en este desafío. Un equipo de desarrollo indexó
sus runbooks, notas de reuniones y documentación de API en borrador en un almacén
vectorial que es accesible desde la aplicación sin ninguna autenticación:

`GET /rag/search`

El endpoint acepta un parámetro de consulta opcional `q` y devuelve los fragmentos
de texto indexados que el almacén considera relevantes para la consulta. Sin una
consulta devuelve todos los fragmentos indexados. Uno de los fragmentos fue
indexado mientras aún contenía una nota en borrador con la credencial del trabajo
nocturno de embeddings.

Recorre los fragmentos indexados a través del endpoint y envía la credencial que
encuentres en ellos.

[NOTE]
====
El endpoint y todas las credenciales de este desafío son ficticios. No se envía
nada por la red: la aplicación convierte los fragmentos en embeddings y los indexa
en su propio índice vectorial en memoria al iniciarse, y solo busca en él.
====
32 changes: 32 additions & 0 deletions src/main/resources/explanations/challenge76_fr.adoc
Original file line number Diff line number Diff line change
@@ -0,0 +1,32 @@
=== Challenge 76 : Trouvez le secret indexé dans le magasin vectoriel RAG

Les pipelines de génération augmentée par récupération (RAG) répondent aux
questions en recherchant d'abord les passages pertinents dans un magasin vectoriel,
puis en alimentant ces passages avec un grand modèle de langage. Pour construire ce
magasin, les documents sont découpés en fragments, convertis en embeddings et
indexés. Les fragments eux-mêmes restent lisibles : un point d'accès de recherche
renvoie le texte original de chaque fragment jugé pertinent.

C'est exactement ce qui s'est mal passé dans ce challenge. Une équipe de
développement a indexé ses runbooks, comptes rendus de réunion et documentation API
de brouillon dans un magasin vectoriel accessible depuis l'application sans aucune
authentification :

`GET /rag/search`

Le point d'accès accepte un paramètre de requête optionnel `q` et renvoie les
fragments de texte indexés que le magasin juge pertinents par rapport à elle. Sans
requête, il renvoie tous les fragments indexés. L'un des fragments a été indexé
alors qu'il contenait encore une note de brouillon avec l'identifiant du travail
d'embedding nocturne.

Parcourez les fragments indexés via le point d'accès et soumettez l'identifiant que
vous y trouvez.

[NOTE]
====
Le point d'accès et tous les identifiants de ce challenge sont fictifs. Rien n'est
envoyé sur le réseau : l'application convertit les fragments en embeddings et les
indexe dans son propre index vectoriel en mémoire au démarrage, puis ne fait que
chercher dans celui-ci.
====
14 changes: 14 additions & 0 deletions src/main/resources/explanations/challenge76_hint.adoc
Original file line number Diff line number Diff line change
@@ -0,0 +1,14 @@
You can solve this challenge using the following steps:

1. Open the unauthenticated search endpoint of the in-memory vector store:
- `http://localhost:8080/rag/search`
- It returns all four indexed text chunks as JSON.

2. Read the chunks and look for the one that mentions a credential:
- The chunk with the "Draft API note" contains the credential of the nightly
embedding job.

3. Submit that credential (everything after "uses the credential ") as the answer.

You can also narrow the search, for example with `q=credential`, but the secret is
only exposed because the chunk was indexed before the draft note was scrubbed.
Loading
Loading