Introduction
I recently integrated the @sentry/cloudflare SDK into a project running on Cloudflare Workers + Workflows. Within days, the Sentry dashboard was flooded with a large number of unresolved issues and hundreds of events.
A closer look revealed that the vast majority were transient failures from external APIs — errors that resolved themselves after retries:
| Error Type | Proportion | Typical Error Message |
|---|---|---|
| 500 Internal Server Error | Highest | 500 Internal Server Error |
| JSON Format Error | High | Invalid JSON |
| Request Timeout | Medium | Request timed out |
| Network Disconnection | Low | Network connection lost |
The core contradiction: Most of these errors are automatically recovered by Cloudflare Workflow's retry mechanism, yet every intermediate failure appears in Sentry. The alert channel gets drowned in noise, making real failures invisible.
Cloudflare Workflow's Retry Mechanism
Cloudflare Workflows splits tasks into retryable atomic steps via step.do(). Each step can be independently configured with a retry policy:
const result = await step.do(
"Call External API",
{ retries: { limit: 3, delay: "2 seconds", backoff: "exponential" } },
async () => {
// Call external API, may fail due to transient errors...
}
)
A typical Workflow contains multiple steps forming a pipeline:
graph LR A[Step 1<br/>Data Preparation] --> B[Step 2<br/>API Call A] B --> C[Step 3<br/>API Call B] C --> D[Step 4<br/>Write Results]
The problem surfaces in successful-retry scenarios. Consider a step with three attempts:
sequenceDiagram participant W as Workflow Engine participant S as step.do() participant API as External API participant Sentry as Sentry W->>S: Execute step S->>API: Attempt 1 API-->>S: 500 Error S->>Sentry: captureException ① Note over S: throw → engine catches, waits for retry W->>S: Retry step S->>API: Attempt 2 API-->>S: 500 Error S->>Sentry: captureException ② Note over S: throw → engine catches, waits for retry W->>S: Retry step S->>API: Attempt 3 API-->>S: 200 OK Note over S: Success, proceed to next step Note over Sentry: Final result: SUCCESS<br/>Yet Sentry has 2 false errors
The 3rd attempt succeeds, the entire Workflow completes normally. But Sentry already has 2 meaningless error events. When Workflows run frequently, noise accumulates to hundreds or thousands of events rapidly.
SDK Source Code: Where the Noise Comes From
Why does every step retry throw get reported to Sentry? The answer lies in @sentry/cloudflare's instrumentWorkflowWithSentry implementation.
The WrappedWorkflowStep.do() Catch Block
The SDK wraps step.do() and automatically captures exceptions when the callback throws. The critical code is at line 80-81 of workflows.js:
// @sentry/cloudflare SDK source (simplified)
async do(name, configOrCallback, callback) {
// ... parameter handling omitted
return this._step.do(name, config, async () => {
return Sentry.startSpan(/*...*/, async (span) => {
try {
return await callback()
} catch (error) {
// Line 80-81: reports on EVERY throw!
captureException(error, {
mechanism: { handled: true, type: 'auto.faas.cloudflare.workflow' }
})
throw error // Re-throw for Workflow engine retry
}
})
})
}
The key: captureException executes before throw. Regardless of whether this error will be successfully retried later, the SDK reports it immediately.
No Catch at the run() Level
Now look at the run() wrapper:
// SDK's run() wrapper (simplified)
async run(event, step) {
return Sentry.withIsolationScope(async () => {
try {
return await originalRun.call(this, event, wrappedStep)
} finally {
// Only finally, no catch
// → If all retries exhaust and the workflow truly fails,
// nobody reports the final error
}
})
}
This creates two problems:
- Intermediate failures are over-reported — every step throw gets reported, even if a subsequent retry succeeds
- Final failures are not reported —
run()only hastry-finally, nocatch
The following flowchart compares the automatic capture path with the ideal path:
flowchart TB
subgraph problem["Current: SDK Auto-Capture"]
direction TB
P1[step callback throws] --> P2[SDK catch block]
P2 --> P3["captureException<br/>mechanism: auto.faas.cloudflare.workflow"]
P3 --> P4[throw → engine retries]
P4 -->|retry succeeds| P5["False error in Sentry ❌"]
P4 -->|retries exhausted| P6["run() finally<br/>no catch → no report ❌"]
end
subgraph solution["Goal: Only Report Final Failures"]
direction TB
S1[step callback throws] --> S2[SDK catch block]
S2 --> S3["captureException → beforeSend filters → discarded"]
S3 --> S4[throw → engine retries]
S4 -->|retry succeeds| S5["No noise in Sentry ✅"]
S4 -->|retries exhausted| S6["run() catch<br/>manual captureException ✅"]
end
GitHub Issue #17421
This isn't a problem unique to any single project. Sentry JavaScript SDK's Issue #17421 discusses the same concern.
A Sentry team member gave a definitive response:
The SDK cannot detect the retry count. Cloudflare Workflow step executions don't share state between retries. Unless Cloudflare exposes retry count metadata in the API in the future, SDK-side improvements are not possible.
In other words, this won't be fixed at the SDK level — consumers need to handle it themselves.
mechanism.type: The Key to Differentiation
Since the SDK reports on every step throw, we need a way to distinguish automatic captures from manual reports. The answer is mechanism.type.
What Is mechanism
Sentry's Exception Interface defines a mechanism field that records metadata about how an exception was captured. The type is a string identifying the capture source.
The naming follows Sentry's Trace Origin RFC convention: auto.<category>.<integration>.<part>.
Comparing mechanism.type Across Sources
| Capture Method | mechanism.type | Meaning |
|---|---|---|
| SDK step auto-capture | auto.faas.cloudflare.workflow |
Reported when Workflow step callback throws |
Manual Sentry.captureException() |
generic |
Developer-initiated report (default) |
| SDK queue auto-capture | auto.faas.cloudflare.queue |
Reported when Queue consumer throws |
This is exactly what we need: SDK auto-captured step errors carry auto.faas.cloudflare.workflow, while manually reported final failures from run() carry generic.
Stability Assessment
The specific values of mechanism.type are not part of Sentry's stable public API. However, the actual risk of change is low:
- The values follow a well-defined naming convention (Trace Origin RFC)
- Historical changes have been minimal (from early
'cloudflare'to the current format, PR #17582) - Even if adjusted in the future, the
auto.faas.cloudflareprefix will likely be preserved - Checking the changelog during SDK upgrades provides adequate protection
Solution Design and Trade-offs
After understanding the SDK behavior, three approaches are possible:
| Approach | Method | Pros | Cons |
|---|---|---|---|
Remove instrumentWorkflowWithSentry |
Skip SDK wrapping, initialize Sentry manually | No noise | Loses automatic tracing, spans, and context injection |
| Swallow exceptions in step callbacks | Catch without re-throwing | No reporting | Breaks Workflow retry mechanism — steps won't retry |
| beforeSend filter + manual run() reporting | Filter auto-captures, manually report final failures | Retains all SDK capabilities, only reports real failures | Depends on mechanism.type value stability |
The third approach is optimal: keep instrumentWorkflowWithSentry for its tracing and context capabilities, use beforeSend to precisely filter noise, and manually report genuine final failures at the run() level.
Implementation
beforeSend Filter
Add a beforeSend hook to the Sentry configuration:
import * as Sentry from "@sentry/cloudflare"
function getSentryOptions(env: Env): Sentry.CloudflareOptions {
return {
dsn: env.SENTRY_DSN,
environment: env.SENTRY_ENVIRONMENT,
tracesSampleRate: 0.1,
// ... other config
beforeSend(event) {
// Filter intermediate retry errors auto-captured by Workflow steps
const isWorkflowAutoCapture = event.exception?.values?.some(
(v) => v.mechanism?.type === "auto.faas.cloudflare.workflow"
)
if (isWorkflowAutoCapture) return null // Discard
return event // Keep
},
}
}
The logic is straightforward: iterate through the event's exception.values, and if any has a mechanism.type of auto.faas.cloudflare.workflow, return null to discard the entire event. All other events (manual reports, queue auto-captures, etc.) pass through normally.
Top-Level try-catch in Workflows
The beforeSend filter solves the noise problem but introduces a gap: if all retries are exhausted and the Workflow truly fails, who reports it?
The answer is adding a top-level try-catch in each Workflow's run() method, manually calling Sentry.captureException():
export class MyWorkflow extends WorkflowEntrypoint<Env, Payload> {
async run(event: WorkflowEvent<Payload>, step: WorkflowStep): Promise<void> {
try {
// step1: Data preparation
const data = await step.do("Load data", async () => { /* ... */ })
// step2: Call external API (configured with 3 retries)
const result = await step.do("Call API",
{ retries: { limit: 3, delay: "2 seconds", backoff: "exponential" } },
async () => { /* ... */ }
)
// step3: Write results
await step.do("Write results", async () => { /* ... */ })
} catch (error) {
// Final failure after all retries exhausted → manual report
// mechanism.type defaults to "generic", won't be filtered by beforeSend
Sentry.captureException(error)
throw error // Re-throw so Workflow is marked as failed
}
}
}
All Workflows follow the same pattern: a top-level try-catch captures final failures, reports manually, then re-throws.
Unit Tests
Comprehensive unit tests were written for the beforeSend logic, covering key scenarios:
describe("beforeSend filters Workflow step auto-captured errors", () => {
it("filters events with mechanism.type auto.faas.cloudflare.workflow", () => {
const event = {
exception: {
values: [{
type: "Error",
value: "some transient error",
mechanism: { type: "auto.faas.cloudflare.workflow", handled: true },
}],
},
}
expect(beforeSend(event, {})).toBeNull() // Discarded
})
it("keeps manual captureException events (mechanism.type generic)", () => {
const event = {
exception: {
values: [{
type: "Error",
value: "final failure",
mechanism: { type: "generic", handled: true },
}],
},
}
expect(beforeSend(event, {})).toBe(event) // Kept
})
it("keeps other Cloudflare auto-captured events like queue", () => {
const event = {
exception: {
values: [{
mechanism: { type: "auto.faas.cloudflare.queue", handled: false },
}],
},
}
expect(beforeSend(event, {})).toBe(event) // Kept
})
it("filters when any exception value matches", () => {
const event = {
exception: {
values: [
{ mechanism: { type: "generic" } },
{ mechanism: { type: "auto.faas.cloudflare.workflow", handled: true } },
],
},
}
expect(beforeSend(event, {})).toBeNull() // Discarded
})
})
Tests cover: filtering workflow auto-captures, keeping manual reports, keeping queue auto-captures, keeping events without mechanism, keeping events without exception, and multi-exception-value scenarios.
Conclusion
Result
After deployment, only genuinely concerning final failures appear in Sentry — errors that persist after all retries are exhausted. Transient failures (API 500s, timeouts, network disconnections) no longer generate noise.
Key Takeaways
-
mechanism.typeis the key to distinguishing auto-captures from manual reports. SDK auto-captured exceptions carryauto.faas.cloudflare.workflow, while manualcaptureException()defaults togeneric. -
instrumentWorkflowWithSentryremains valuable. It provides automatic tracing, span creation, and context injection. We only need to filter its error reporting, not remove the entire wrapper. -
beforeSendis Sentry's most flexible client-side filtering mechanism. It intercepts events before sending and can filter, modify, or discard based on any event property.