AI-Generated Terraform

The Hidden Risk in AI-Generated Terraform

Ask ChatGPT or Claude for a Terraform module to stand up a three-tier AWS deployment, and it will hand you something that looks production-ready in under a minute. That speed is the appeal, and it’s why a fast-growing share of infrastructure teams now draft code this way. It’s also the trap. 

The model has no idea what your compliance posture requires, what your last incident was, or which security group should never be opened to the entire internet. It knows what a plausible answer looks like, and it optimizes for that, not for what your auditor will ask about in six months.

That gap shows up in the numbers, too. Veracode’s 2026 GenAI Code Security Report tested how leading language models handle real coding tasks and found the average security pass rate has stalled at 56%, flat versus a year earlier. About 44% of the tasks in the study introduced an exploitable weakness. Even the best-performing model in the report still failed close to one in three security checks. 

Spacelift’s own research on infrastructure and platform teams found a similar pattern on the buyer side: 78% of organizations already use AI to generate infrastructure as code without review, and 93% have experienced at least one AI-related infrastructure incident.

For a marketing page, a bug like that is embarrassing. For infrastructure code with the power to create, modify, or destroy cloud resources the moment it’s applied, it’s an open door.

The review step nobody skips, until they do

“Every AI-generated change should go through review before merge” is the standard advice, and it’s sound. It’s also the first casualty of a deadline. With a plan output that reads clean and a deploy window closing, review turns into a rubber stamp or gets skipped outright. 

Nobody decided to abandon governance. Nothing in the workflow made skipping it harder than approving it, and a model that writes confidently, right or wrong, makes that rubber stamp easier to justify.

What governed AI provisioning actually looks like

The fix isn’t asking humans to review AI output more carefully. It’s removing the moment where a bad decision can slip through unchecked. That’s precisely the gap Spacelift Intelligence is built to close: an AI layer covering the whole lifecycle of a change instead of only the moment a developer types a request.

Intent is the deployment piece of it: a developer describes what they need, a QA environment, a demo, a short-lived experiment, in plain English, and the AI provisions it directly through the cloud provider’s API instead of generating a Terraform file for someone to review later. Intelligence doesn’t stop once something is running, though. It can explain what changed and why in plain language, flag risky patterns before a deployment goes out, and turn policy as code itself into something you describe in a sentence instead of writing from scratch. 

The part that actually matters ties it all together: whether a change started as an Intent request or a standard Terraform run, it passes through the same policy engine. A request that violates policy gets flagged or blocked before it provisions anything, not caught in a retrospective audit three weeks later. Governance didn’t slow down to make room for AI. It became automatic, the only version of it that survives a deadline.

That same logic applies to conventional infrastructure as code, AI-written or not. Terraform guardrails, enforced through policy as code, check a plan before it ever reaches an apply and block the public bucket, the unencrypted database, or the wide-open security group automatically, regardless of who or what wrote the configuration. 

Requirements like SOC 2, HIPAA, and PCI DSS stop depending on a reviewer’s memory and become something the pipeline enforces on its own. Drift detection covers the other side of the problem no code review could ever catch: the change made outside the pipeline entirely, after deployment.

The next phase of this is already underway. Instead of a developer copying AI output into a terminal, AI agents are starting to call infrastructure tools directly, through interfaces such as the Model Context Protocol. They pull live provider documentation and trigger runs with far less human intervention between prompt and production. That’s a bigger lever and a bigger risk delivered in the same motion. It needs the same answer as Intent: read operations that run automatically, destructive ones that stay disabled until a human turns them on, and access scoped to the smallest project it touches rather than an entire account.

None of this is an argument against using AI to write or provision infrastructure. It’s an argument against mistaking the model’s output for the finish line rather than the first draft. The teams getting real value out of AI-assisted infrastructure aren’t the ones typing faster prompts. They’re the ones who paired that speed with a policy engine that never gets tired, never has a deadline, and never rubber-stamps anything by accident.

Leave a Comment

Your email address will not be published. Required fields are marked *

Scroll to Top