Skip to content

[BUG] [v0.0.7] CreateAgent tool writes .toml files into .cortex/agents, so created agents are not discoverable by the standard loader #53529

Description

@DDDDDGCSM

Project

cortex

Description

The CreateAgent tool reports success and writes a file under .cortex/agents/, but it writes <name>.toml. The standard custom-agent loader only discovers Markdown agent files, so the created agent is immediately invisible to normal agent discovery and cannot be delegated to through the standard loader path.

Error Message

No explicit error. The tool returns success, but the created agent does not appear in the standard loader results.

Debug Logs

Focused test:
test_create_agent_project_output_is_not_discoverable_by_standard_loader

Observed file created by CreateAgent:
.cortex/agents/process-and-transform-data-files.toml

Observed loader search paths:
- .agents
- .agent
- .cortex/agents
- ~/.cortex/agents
- ~/Library/Application Support/cortex/agents

Observed loader result:
loaded_names=[]

System Information

cortex source tree v0.0.7
macOS arm64

Screenshots

https://raw.githubusercontent.com/DDDDDGCSM/solid-carnival/master/evidence/cortex-create-agent-toml-proof.png

Steps to Reproduce

  1. Invoke the CreateAgent tool with a valid project-scoped description.
  2. Let it create the agent under .cortex/agents/.
  3. Run the standard custom-agent loader over the same project root.
  4. Observe that the loader returns no agent with that name.

Expected Behavior

A newly created custom agent should be written in the same file format that the standard loader understands, so the created agent is immediately discoverable and usable.

Actual Behavior

CreateAgent writes a TOML file into .cortex/agents/, but the standard loader only discovers Markdown agent files, so the created agent is effectively orphaned.

Additional Context

Code evidence in v0.0.7:

  • src/cortex-engine/src/tools/handlers/create_agent.rs writes agents_dir.join(format!("{}.toml", agent_name))
  • src/cortex-agents/src/custom/loader.rs loads standard agent directories and reads Markdown agent files (*.md)

This is independent from the app-server .factory/agents path mismatch: here the directory is nominally correct, but the file format is incompatible with the loader.

Activity

  1. atlas-cortex commented on May 12, 2026

    @atlas-cortex

    ✅ Validation Summary — Issue #53529

    All validation checks passed. 1/17 code rules failed | 1 PENALIZE | 14 LLM instructions active

    Confidence Score: 4/5

    • All validation checks passed. This submission meets the bounty requirements.

    Validation Checks

    Check Status Detail
    Media Evidence ✅ Pass Evidence found and accessible
    Spam Detection ✅ Pass Score: 23% — no spam signals
    Duplicate Detection ✅ Pass No significant overlap with existing issues
    Code Verification ✅ Pass Bug confirmed in source code (confidence: 100%) — The code confirms that CreateAgentHandler writes .toml files, while `CustomA
    Edit History ✅ Pass No suspicious edits
    LLM Evaluation ✅ Pass All validation checks passed. 1/17 code rules failed

    Validation Pipeline

    %%{init: {'theme': 'neutral'}}%%
    flowchart TD
        A["Issue #53529"] --> B{"Media Check"}
        B -->|"✅ Found"| C{"Spam Check"}
        C -->|"✅ Clean"| D{"Duplicate Check"}
        D -->|"✅ Unique"| CV{"Code Verify"}
        CV -->|"✅ Confirmed"| E{"Edit History"}
        E -->|"✅ Normal"| F{"LLM Evaluation"}
        F -->|"✅ Valid"| G["✅ Approved"]
        style G fill:#51cf66,color:#fff
    
    Loading

    Raw Evidence

    {
      "media": {
        "hasMedia": true,
        "accessible": true,
        "urls": [
          "https://raw.githubusercontent.com/DDDDDGCSM/solid-carnival/master/evidence/cortex-create-agent-toml-proof.png"
        ],
        "evidence": [
          "Found 1 media URL(s) via markdown parsing",
          "1 URL(s) accessible"
        ]
      },
      "spam": {
        "overallScore": 0.23262711864406782,
        "details": "template=0.39 (1 recent issues compared); burst=0.25 (2 issues in 2h window); parity=0.00; overall=0.23 (threshold=0.7)"
      },
      "duplicate": {
        "isDuplicate": false,
        "similarity": 0.6532722120922858,
        "topSimilar": [
          {
            "issueNumber": 53528,
            "title": "[BUG] [v0.0.7] app-server /agents API writes to .factory/agents, so standard loaders never discover created agents",
            "similarity": 0.6532722120922858
          },
          {
            "issueNumber": 52221,
            "title": "[BUG] [v0.0.7] cortex agent create --non-interactive ignores CORTEX_HOME and writes to the default personal agents directory",
            "similarity": 0.5522942617117701
          },
          {
            "issueNumber": 52368,
            "title": "[BUG] [v0.0.7] cortex agent create writes agent .md files with world-readable permissions (0664) — system prompts and tool configurations exposed",
            "similarity": 0.5373983517191413
          },
          {
            "issueNumber": 52261,
            "title": "[BUG] [v0.0.7] cortex agent copy ignores CORTEX_HOME and writes the copied agent to the default personal agents directory",
            "similarity": 0.5303648248438788
          },
          {
            "issueNumber": 51577,
            "title": "cortex agent list/show ignore CORTEX_HOME — always scan ~/.cortex/agents",
            "similarity": 0.5296367116825444
          }
        ]
      },
      "editHistory": {
        "fraudScore": 0,
        "editCount": 0,
        "details": "No edits detected"
      },
      "codeVerification": {
        "plausible": true,
        "confidence": 1,
        "reasoning": "The code confirms that `CreateAgentHandler` writes `.toml` files, while `CustomAgentLoader` only searches for and processes `.md` files. This mismatch makes agents created by the tool undiscoverable by the loader.",
        "codeEvidence": "src/cortex-engine/src/tools/handlers/create_agent.rs writes agents with a .toml extension and TOML content. src/cortex-agents/src/custom/loader.rs only loads agents with a .md extension and expects YAML frontmatter.",
        "screenshotValid": true,
        "screenshotReasoning": "The screenshot analysis returned valid."
      },
      "rules": {
        "codeResults": {
          "passed": [
            {
              "ruleId": "content.no-profanity",
              "category": "content",
              "severity": "flag",
              "passed": true,
              "message": "Issue should not contain excessive profanity or abusive language",
              "weight": 1
            },
            {
              "ruleId": "content.reasonable-length",
              "category": "content",
              "severity": "flag",
              "passed": true,
              "message": "Issue body should not exceed 15000 characters (possible spam dump)",
              "weight": 1
            },
            {
              "ruleId": "content.has-context",
              "category": "content",
              "severity": "penalize",
              "passed": true,
              "message": "Issue should mention the affected page, component, or URL",
              "weight": 0.2
            },
            {
              "ruleId": "media.require-evidence",
              "category": "media",
              "severity": "require",
              "passed": true,
              "message": "Issue must include at least one screenshot or video URL",
              "weight": 1
            },
            {
              "ruleId": "media.must-be-accessible",
              "category": "media",
              "severity": "require",
              "passed": true,
              "message": "All media URLs must be publicly accessible (HTTP 200)",
              "weight": 1
            },
            {
              "ruleId": "media.no-placeholder-urls",
              "category": "media",
              "severity": "reject",
              "passed": true,
              "message": "Media URLs should not be placeholder or example URLs",
              "weight": 1
            },
            {
              "ruleId": "scoring.suspicious-edits",
              "category": "scoring",
              "severity": "penalize",
              "passed": true,
              "message": "Penalize issues with concerning edit history (fraud score > 0.3)",
              "weight": 0.25
            },
            {
              "ruleId": "spam.high-score-reject",
              "category": "spam",
              "severity": "reject",
              "passed": true,
              "message": "Reject issues with spam score above 0.85",
              "weight": 1
            },
            {
              "ruleId": "spam.generic-title",
              "category": "spam",
              "severity": "penalize",
              "passed": true,
              "message": "Title should not be a generic/template title",
              "weight": 0.4
            },
            {
              "ruleId": "spam.body-is-title-repeat",
              "category": "spam",
              "severity": "penalize",
              "passed": true,
              "message": "Body should not be a simple repetition of the title",
              "weight": 0.5
            },
            {
              "ruleId": "spam.no-ai-filler",
              "category": "spam",
              "severity": "flag",
              "passed": true,
              "message": "Body should not contain obvious AI-generated filler phrases",
              "weight": 1
            },
            {
              "ruleId": "validity.min-body-length",
              "category": "validity",
              "severity": "reject",
              "passed": true,
              "message": "Issue body must be at least 50 characters",
              "weight": 1
            },
            {
              "ruleId": "validity.min-title-length",
              "category": "validity",
              "severity": "reject",
              "passed": true,
              "message": "Issue title must be at least 10 characters",
              "weight": 1
            },
            {
              "ruleId": "validity.no-empty-body",
              "category": "validity",
              "severity": "reject",
              "passed": true,
              "message": "Issue body must not be empty or only whitespace",
              "weight": 1
            },
            {
              "ruleId": "validity.has-steps-or-description",
              "category": "validity",
              "severity": "penalize",
              "passed": true,
              "message": "Issue body should contain structured content (steps, expected/actual behavior)",
              "weight": 0.3
            },
            {
              "ruleId": "validity.not-a-feature-request",
              "category": "validity",
              "severity": "flag",
              "passed": true,
              "message": "Issue should describe a bug, not a feature request",
              "weight": 1
            }
          ],
          "failed": [
            {
              "ruleId": "scoring.duplicate-threshold",
              "category": "scoring",
              "severity": "penalize",
              "passed": false,
              "message": "Issue has moderate similarity to existing issues, suggesting partial overlap.",
              "weight": 0.3
            }
          ]
        },
        "llmInstructions": [
          "Always prioritize concrete evidence (screenshots, videos, URLs) over the quality of the written description. A poorly written report with clear visual proof of a bug is VALID. A beautifully written report with no evidence is INVALID.",
          "A valid bug report must contain enough information for an engineer to reproduce the issue. If the steps to reproduce are missing or impossibly vague (\"it just broke\"), the issue is INVALID.",
          "You MUST call the deliver_verdict function exactly once. Do not output a JSON object manually. Do not output your verdict in plain text. Use the tool.",
          "Be highly suspicious of submissions that follow a rigid template: same structure, same phrasing, with only the page name or URL swapped. These are template-farmed and should be INVALID. Look for: identical sentence structure, placeholder-like descriptions, formulaic \"Steps to Reproduce\".",
          "When similar issues exist, the OLDER issue (lower issue number) always takes precedence. The newer submission is the duplicate, never the other way around.",
          "Confidence scores must be calibrated: use 0.9+ only when evidence is unambiguous and complete. Use 0.5-0.7 when the issue is borderline. Use below 0.5 when you are genuinely uncertain. Never default to 0.5 — commit to a direction.",
          "Do not give the benefit of the doubt. Missing evidence means INVALID. Unclear reproduction steps means INVALID. Inaccessible media means INVALID. The burden of proof is on the submitter.",
          "In the reasoning field, explain your step-by-step analysis BEFORE stating your conclusion. Cover: evidence quality, clarity, spam signals, and duplicate overlap in that order.",
          "Never mention internal scoring thresholds, rule IDs, detection heuristics, or system architecture in the recap field. The recap is public-facing. Keep reasoning technical but recap user-friendly.",
          "If the issue body reads like AI-generated boilerplate (phrases like \"It is important to note\", \"Furthermore\", \"In conclusion\", excessive politeness, no specific technical details), treat it as strong evidence of spam. AI-generated issues with no real bug details are INVALID.",
          "If the pre-computed media check says media URLs exist, verify the description matches what a screenshot would show. If the description talks about a login page but the context suggests the screenshot is of something unrelated, flag this as suspicious.",
          "Never soften a verdict out of sympathy. \"I can see you tried hard but...\" is not acceptable. If the submission does not meet criteria, it is INVALID regardless of apparent effort.",
          "Write your reasoning and recap in a professional, neutral tone. Do not be sarcastic, condescending, or emotional. State facts.",
          "The recap field must be 2-3 sentences maximum. It will be posted as a public comment. No bullet points in recap — use flowing prose."
        ],
        "totalCodeRules": 17,
        "totalLLMRules": 14,
        "hasReject": false,
        "hasFailed": true,
        "penaltyScore": 0.3,
        "summary": "1/17 code rules failed | 1 PENALIZE | 14 LLM instructions active"
      }
    }

    Validated by Atlas • 2026-05-12

  2. SlaVKsVolks commented on May 12, 2026

    @SlaVKsVolks

    Opened upstream Cortex fix: CortexLM/cortex#273

    What changed:

    • CreateAgent now writes custom agents as Markdown files with YAML frontmatter under .cortex/agents, matching the standard CustomAgentLoader format.
    • The generated agent remains in the standard project agent directory but is now immediately discoverable by the loader.
    • Added regression coverage that creates a project agent and verifies the standard loader can discover it.

    Verification:

    • cargo fmt --check --package cortex-engine
    • cargo test -p cortex-engine test_create_agent -- --nocapture
    • cargo check -p cortex-engine
    • cargo build -p cortex-engine
    • git diff --check

    Note: strict clippy with -D warnings is currently blocked by pre-existing unrelated warnings in cortex-windows-sandbox and cortex-engine Windows sandbox/config code.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    ideIssues related to IDEvalidValid issue

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions