Claude Code in Headless Non-Interactive Mode
Headless mode unlocks automation by moving human judgment into configuration ahead of time.

Claude Code runs in two fundamentally different ways: as an interactive REPL where a developer types prompts and reviews changes in real time, or as a headless process triggered by the -p flag, where a prompt goes in, Claude does the work, the result prints to stdout, and the process exits. The second mode is what makes Claude Code usable inside CI/CD pipelines, cron jobs, and multi-agent systems. This piece covers how headless mode actually works, how to configure it so it holds up under production load, and where the risks sit when nobody is watching the terminal.
How headless mode differs from interactive use
A typical headless invocation looks like cat build-error.txt | claude -p "Explain the root cause…". The prompt arrives piped on stdin, assembled from the output of another command, the natural shape for a pipeline step. Piped input has a size cap, and oversized input gets rejected outright rather than quietly cut down to fit, so a script feeding Claude a large log file needs to check size before it pipes anything in.
The prompt can also be passed as a plain positional argument, but the stdin pattern is what matters for automation, because it lets a prompt get built dynamically from whatever other tools in the pipeline produce. That's the surface-level mechanic. The deeper shift is in what disappears when the human leaves the loop. In an interactive session, a developer reads Claude's reasoning as it happens, answers permission prompts one at a time, and decides when a task is actually finished. None of that judgment exists by default in headless mode. Every decision a human would normally make on the fly, whether to grant a permission, whether the output looks right, whether the task has run long enough, has to be written into flags, config files, and surrounding pipeline logic ahead of time. Headless mode doesn't remove those decisions; it moves them earlier, into the setup, where they need to be right before the job ever runs.
Why the interactive loop breaks down at automation scale
An interactive session depends on a person being present to answer prompts, a poor fit for repeatable or scheduled work. A pipeline running a deployment can't pause and wait for someone to approve a step. A cron job firing at 3 a.m. has no one logged in to click "yes." Any workflow that needs to run the same way, on a schedule, across many repositories, needs a mode that doesn't stall waiting for a human.
The default permission system makes this more pressing, not less. Claude Code asks for approval before running commands or editing files, which sounds like a safeguard, and in principle it is. But Anthropic reports that users approve 93% of these prompts anyway, so the friction stays in place while the protection it's supposed to give barely works, because most people click through without scrutinizing each one.
Headless mode earns its place by unlocking a specific category of work: analyzing test failures inside a deployment pipeline, updating documentation the moment a schema changes, refactoring deprecated API calls across a fleet of microservices. These are tasks where the entire point is that no one has to sit and supervise them.
That said, headless mode has a real ceiling. A single -p call can't do what an interactive session does when it reads files, asks a clarifying question, and changes direction mid-task based on the answer. Headless mode fits well-defined work, where the input is clear and the output is clear. Anything exploratory, anything that needs course correction as it goes, still belongs in the REPL.
Configuring output format so downstream scripts can parse results reliably
Once a task is automated, the first real configuration decision is how Claude's output gets formatted, because everything downstream, error handling, cost tracking, aggregating results across agents, depends on whether that output can be parsed reliably.
The --output-format flag gives three choices. The default, text, returns a human-readable response, which works fine for a quick one-off script or for piping Claude's answer straight to a person. json returns a single JSON object containing the response text, token usage, cost, and a session ID, and it's the right choice for any programmatic caller, for cost tracking, and for audit logs. stream-json returns newline-delimited JSON, one event per message, fitting a UI that needs to show live progress as Claude works.
The JSON object carries everything a wrapper script actually needs: a type, a subtype, the result itself, a session_id, total_cost_usd, and a usage block that breaks down input tokens, output tokens, cache read tokens, and cache creation tokens. With that structure, a script knows what a given run cost and which session ID to resume from if it needs to pick the conversation back up later.
Permission denials are the failure mode worth taking seriously here. A production pattern worth adopting is to pipe the JSON output into jq and check the permission_denials array; if that array is non-empty, the pipeline should treat the run as a failure, even if Claude exits with status code 0. A run that gets blocked from editing a file or running a command can still exit cleanly, and a pipeline that only checks the exit code will wave that failure straight through as a success. That's the single most costly oversight in a headless setup: silent, false-positive success.
When to disable session persistence
Claude Code saves a session transcript, the conversation history and the results of any tool calls, to disk by default. That transcript isn't loaded automatically on the next run, though. Reloading it requires explicitly passing --continue or --resume. For a lot of automated tasks, this default is actually fine precisely because it's opt-in: nothing carries over unless asked for.
But the opt-in nature of persistence is also why it has to be managed deliberately. Session persistence makes sense for tasks that build on themselves over time, documentation that gets refined across several days of runs, for instance. It's the wrong setting for build-time code generation, where a CI job needs to produce identical output every single time it runs, regardless of what happened in any prior run.
The failure mode here is quiet but costly: a headless agent that keeps picking up stale context across separate runs starts producing results that drift and become inconsistent, and tracking down why requires piecing back together an entire session history that nobody was watching as it accumulated, so when resuming a session really is the intent, the session_id returned in the JSON output is the handle for it.
Permission modes for unattended runs and the risk spectrum between them
Permissions are the single highest-stakes configuration decision in headless automation. Getting them wrong in the permissive direction means an unattended run can push to a branch, delete files, or call an external API with no one watching it happen.
Three approaches sit along a real spectrum, each with a legitimate use case. Allowlisting specific tools through a flag like --allowedTools "Read,Grep" sits at the cautious end: medium risk, with the operator defining exactly which tools Claude can touch. A code review agent, for example, should only ever get read access, never Edit or Bash. At the other end, --dangerously-skip-permissions turns off every prompt, so Claude can act without asking. That setting makes sense only inside infrastructure that bounds the blast radius on its own, an ephemeral VM, a disposable container, a sandboxed git worktree, where whatever Claude does can't escape the box it's running in.
A proof-of-concept from Adversa shows how far that flag opens the door: with --dangerously-skip-permissions active, a sequence of 50 no-op subcommands ran, followed by a curl call that executed with no prompt whatsoever. The flag doesn't just lower the approval threshold. It removes the last runtime check entirely, so whatever comes after those no-ops runs unchecked.
The riskiest configuration is -p, --dangerously-skip-permissions, and a long agentic task running together inside an environment with over-scoped credentials and no mapped boundaries on what the agent can reach. In that combination, an agent can act, call out to other systems, and move data before anyone has a chance to review a single step. A middle setting, often called auto mode, cuts risk compared to skipping permissions outright, but it doesn't replace the underlying work of scoping credentials tightly and isolating where the job runs. Whatever permission mode gets used, pair it with --max-turns N, which caps how many agentic turns, tool-use round trips, the agent can take before it's forced to respond. That cap bounds both cost and runtime no matter which permission setting is active, and it belongs in every production headless run.
Authentication in CI and the credential path to use
In headless mode, credentials come from a CI secret set as an environment variable, not from an interactive login. The execution environment needs ANTHROPIC_API_KEY available directly. A developer's logged-in terminal session doesn't help a job running on a build server somewhere else.
Claude Code supports a few authentication paths, and they carry meaningfully different billing and rate-limit behavior. A Max subscription, set up through claude login, stores OAuth tokens and charges a flat monthly fee, with a 5-hour rolling session limit and weekly rate limits on top. It suits predictable workloads that stay comfortably under those caps. An API key, set through the ANTHROPIC_API_KEY environment variable, charges per token with no rate ceiling, which suits bursty workloads, shared team usage, or heavy overnight batch runs. For large organizations already running on AWS or GCP, Claude is also available through Bedrock or Vertex AI, billed through the cloud vendor directly, which fits existing cloud contracts and data-residency requirements without adding a separate billing relationship.
A practical rule of thumb: one to three agents running steadily tend to fit comfortably on a Max subscription. Five or more agents running overnight will hit rate limits, and at that point an API key becomes the right tool for burst capacity. Many teams end up running both: a subscription for day-to-day interactive work, and an API key for the overnight fleet that needs to run past what the subscription caps allow. Switching between the two at runtime is just a matter of setting or unsetting the environment variable. Whichever credential is active at the moment a job runs determines which billing pool gets charged, so a misconfigured environment variable doesn't just break the job, it can also charge the wrong account.
The billing fracture that exposed headless use as a categorically different workload
Anthropic's attempt to restructure billing in June 2026 showed something important about how the company actually treats headless workloads: as a separate category from interactive use, distinct enough to warrant its own pricing logic, regardless of whether that specific policy ultimately survives in its original form.
On June 15, 2026, the exact day the restructure was set to take effect, Anthropic paused it after backlash from the developer community. The Agent SDK, claude -p, and third-party apps kept drawing from regular subscription limits. Then, on October 7, 2026, Max and Team plans gained monthly API credits that cover the Agent SDK and claude -p directly, folding headless usage into the subscription.
Multica, a company built around the tool, illustrates what that near-miss would have cost, since its daemon hardcodes -p as a mandatory flag for every task it runs, because -p is the only way to get the structured JSON stream the daemon needs to parse messages. Had the original restructure gone through without a subscription-compatible headless mode, every task that daemon ran would have drawn on API credits or Agent SDK credits separately from whatever plan the user was already paying for, with no warning inside the product telling them it was about to happen.
The pause buys time, but it doesn't settle anything permanently. Anyone building a production headless workflow now should design it to run correctly on API-key billing as the baseline, treating subscription coverage of -p as a bonus. A pipeline built that way keeps working and keeps its costs predictable no matter which direction Anthropic's policy moves next.
Wiring headless Claude Code into CI/CD pipelines and cron jobs
Every piece covered so far, output format, session state, permissions, authentication, billing, comes together in one fixed shape for a well-built unattended run: a trigger fires, the job authenticates with its provider, Claude Code runs headless inside whatever guardrails were set up in advance, the pipeline parses the structured output it gets back, and only after all of that does it act, whether that means posting a comment, committing code, or opening a pull request.
In practice, that means wrapping the claude -p call inside a shell script, capturing its output for logging, and exiting with a non-zero status whenever the task fails, so the surrounding orchestration system correctly marks it as a failure. The ANTHROPIC_API_KEY should always come in as a CI secret injected into the job's environment, never typed directly into a script file where it can end up committed to a repository by mistake.
None of these pieces work in isolation. Structured JSON output is what lets a script that's only watching exit codes see a failed permission check. Statelessness is what keeps a nightly job from quietly inheriting yesterday's leftover context. Scoped permissions are what keep a bounded task bounded when no one is there to intervene. Put together correctly, they turn Claude Code's headless mode from a convenient trick for one-off scripts into infrastructure that a team can actually build a pipeline on and trust to run the same way every time.