Better Agentic Review with Ephemeral Deploys
Use Miren ephemeral deploys to let agents review PRs
It will come as no surprise to teams doing agentic development that the number of pull requests (PRs) teams submit has gone up precipitously in the last year. Many teams are now solving an agent problem by allowing agents to in turn review those PRs.
When the agent in CI is reviewing a PR, it’s doing so in a constrained environment and likely with a time deadline. If the code being tested is a large application, it’s not uncommon that the app can’t be booted within the CI environment at all. This could be because of its raw CPU and memory footprint, or its service dependencies, or even data regulations. Regardless of the reason, setting up a whole application in CI comes with its own set of issues.
For all those reasons, Miren ephemeral deploys are an ideal way to give the CI agent a full deployment to test against. Miren ephemeral deployments, also called pull request environments, provide the ability to deploy a different version of an application that runs separately from the main application deployment. This different version shares the configuration with the main version, providing access to the same services, tokens, etc.
Combined with CI deployment, we can very easily create an ephemeral deployment from CI such as GitHub Actions to enable the agent to test against.
Let’s go through a whole demo before we dive into the nitty gritty:
GitHub Actions
The whole sample repo with all of this is at github.com/mirendev/agentic-eph.
Let’s break down what this looks like in GitHub Actions.
We begin the workflow with a concurrency group that lets only one run per PR go at a time:
concurrency:
group: pr-preview-${{ github.event.number }}
cancel-in-progress: true
This prevents runs generated by new code being pushed to the PR from overlapping. That overlapping would cause the ephemeral deployments to change and become out of sync with the code the agent is testing.
Next, a preview job deploys the code:
preview:
# Fork PRs don't get secrets, so they can't deploy. Same-repo PRs only.
if: github.event.pull_request.head.repo.full_name == github.repository
runs-on: ubuntu-latest
permissions:
contents: read
id-token: write
outputs:
url: ${{ steps.deploy.outputs.url }}
steps:
- uses: actions/checkout@v4
- id: deploy
uses: mirendev/actions/deploy@main
with:
cluster: ${{ secrets.MIREN_CLUSTER }}
app: brewbar
ephemeral: pr-${{ github.event.number }}
ttl: 24h
The MIREN_CLUSTER secret is created from running miren cluster export-address. The job passes the preview’s URL on as an output, so later jobs can read it as needs.preview.outputs.url.
Ephemeral deployments share their configuration with the base application, including database and other add-ons. For this reason, we almost always suggest that people use ephemeral deployments against a staging version of their application. That prevents the ephemeral deployments from seeing and potentially changing production data. Our PR previews post covers how to set that up.
With the deployment running, you can direct an agent towards it. In our demo, a second job runs Claude Code directly. It waits for preview, needs permission to comment on the PR, and checks out the full git history so the agent can diff against main:
agent-review:
needs: preview
runs-on: ubuntu-latest
permissions:
contents: read
pull-requests: write
id-token: write
steps:
- uses: actions/checkout@v4
with:
fetch-depth: 0 # the agent diffs against origin/main
Now we’re ready for the core of review, invoking the agent:
- name: Agent tests the preview
id: agent
uses: anthropics/claude-code-action@v1
with:
anthropic_api_key: ${{ secrets.ANTHROPIC_API_KEY }}
github_token: ${{ github.token }}
# --json-schema makes the agent's final answer come back as the
# structured_output step output, so it never needs to write a file.
claude_args: |
--model claude-sonnet-5-5
--max-turns 40
--allowedTools "Bash(curl:*),Bash(jq:*),Bash(git diff:*),Bash(git log:*),Bash(git show:*),Read,Glob,Grep"
--json-schema '{"type":"object","additionalProperties":false,"required":["decision","summary","checks"],"properties":{"decision":{"type":"string","enum":["merge","reject"]},"summary":{"type":"string"},"checks":{"type":"array","items":{"type":"object","additionalProperties":false,"required":["name","result","detail"],"properties":{"name":{"type":"string"},"result":{"type":"string","enum":["pass","fail"]},"detail":{"type":"string"}}}}}}'
prompt: |
You are the release gate for Brewbar, a small coffee-ordering web app.
Pull request #${{ github.event.number }}
has been deployed to a temporary preview environment at:
${{ needs.preview.outputs.url }}
Decide whether this pull request is safe to merge into main. Base the
decision on what the running preview actually does, not only on
reading the code.
1. Understand the change.
- `git log --oneline origin/main..HEAD`
- `git diff origin/main...HEAD`
2. Read `docs/api.md`. It is the public API contract. The web page
(`static/index.html`) and a separately released mobile app depend
on it. Adding endpoints or optional fields is fine. Renaming,
removing, or changing the meaning of an existing field or endpoint
is a breaking change, even if the docs were edited to match.
3. Test the live preview with curl. At minimum:
- `GET /healthz` returns 200.
- `GET /api/menu` returns the documented shape.
- `POST /api/orders` happy path: work out the expected
subtotal, discount, tax, and total yourself from the pricing
rules in the contract, and compare them to the response.
- A discount code, and each documented error case with its
documented status code.
- `GET /` serves the page, and every JSON field the page's
JavaScript reads is present in the live responses.
- Whatever new behavior this pull request adds, exercised for
real against the preview.
4. Decide.
- "merge" only if the change works on the preview as described
and nothing in the existing contract broke.
- "reject" if anything is broken, or if you could not verify
the change.
Your shell access is limited to curl, jq, and read-only git
commands. Use the Read tool to read files. Run one command at a
time. Piping curl into jq is fine. Loops, scripts, and `&&`
chains will be blocked. Do the arithmetic yourself.
Return your decision as the structured output:
- decision: "merge" or "reject".
- summary: two or three plain sentences a teammate can read in
ten seconds.
- checks: one entry per thing you tested, with what you sent and
what came back in `detail`.
Do not edit files, and do not merge, comment on, or approve the
pull request yourself. The workflow does that from your decision.
- name: Read the verdict
id: verdict
env:
VERDICT: ${{ steps.agent.outputs.structured_output }}
run: |
mkdir -p .agent
echo "$VERDICT" > .agent/verdict.json
jq -e '(.decision == "merge" or .decision == "reject") and (.checks | type == "array")' .agent/verdict.json
echo "decision=$(jq -r .decision .agent/verdict.json)" >> "$GITHUB_OUTPUT"
jq -r -f .github/agent/report.jq .agent/verdict.json > .agent/report.md
cat .agent/report.md >> "$GITHUB_STEP_SUMMARY" That’s a bit verbose, but it’s designed to keep the agent on rails in this constrained environment. The key though is that you can vary the prompt as much as you need to conform to your own setup. The report.jq script that turns the verdict into a readable report is in the sample repo.
Reporting
Our example also posts the verdict as a comment on the PR. This is great to easily be able to understand its reasoning.
- name: Post the report on the PR
env:
GH_TOKEN: ${{ github.token }}
PR_NUMBER: ${{ github.event.number }}
run: gh pr comment "$PR_NUMBER" --body-file .agent/report.md
Auto-Merging
If you’ve become comfortable with the fidelity of the system, you can even go to the next level and merge the PR based solely on this agent’s decision. I’d only recommend doing this once you’ve used this for a bit and feel the agent’s decisions consistently match your own AND you’re ready to accept a level of risk with automatic, non-human merges.
Part of that risk is the PR itself. The agent reads the PR’s diff and commit messages, which the PR’s author wrote, so a hostile PR could try to talk the agent into saying “merge”. Only auto-merge PRs from contributors you already trust. In our example, the preview job skips PRs from forks, so the agent job and this merge never run for them.
The merge uses a separate MERGE_TOKEN rather than the built-in github.token. GitHub doesn’t start new workflow runs for merges made with the built-in token, so your production deploy on main would never fire.
- name: Merge
if: steps.verdict.outputs.decision == 'merge'
env:
GH_TOKEN: ${{ secrets.MERGE_TOKEN }}
PR_NUMBER: ${{ github.event.number }}
HEAD_SHA: ${{ github.event.pull_request.head.sha }}
HEAD_REF: ${{ github.event.pull_request.head.ref }}
run: |
gh api -X PUT "repos/$GITHUB_REPOSITORY/pulls/$PR_NUMBER/merge" \
-f merge_method=squash -f sha="$HEAD_SHA"
gh api -X DELETE "repos/$GITHUB_REPOSITORY/git/refs/heads/$HEAD_REF"
Conclusion
And there you have it! You can tell everyone on social media that you’ve implemented a robust, fully agentic testing process!