Single-Strategy Deep Dive: Pushing Integration Errors to a DingTalk Robot in Real Time with Qeasy
What This Strategy Solves
In a private-deployment environment where several business systems run side by side, the worst thing for ops is "silent failure": logs pile up inside the platform unread, and the business side only realizes days later that data never arrived. We have run into this exact pattern at a retail customer's site. A few master-data sync strategies failed intermittently; because nothing pushed a notification in real time, the issue sat for two days before anyone noticed, and downstream store systems kept working with stale item master records.
The "DingTalk error-notification" strategy exists for exactly this situation. It does not move business data itself; it acts as a sentinel. Through the Qeasy data-integration platform, it periodically pulls execution records whose status is "error" via a strategy-exception query API, packages them into a structured message, and pushes them to an ops group or a designated DingTalk recipient via a DingTalk custom robot. It is a partner to the usual "item sync" or "sales-order sync" strategies: those strategies move data, this one makes sure that when something breaks, someone actually sees it.
Data Flow and Field Mapping
The whole chain can be summarized as: Qeasy platform (source, provides the error data) → middle assembly (scripts/templates) → DingTalk robot (target, executes the push).
| Key field | Origin | Meaning | Role in this strategy |
|---|---|---|---|
| recentSeconds | Strategy input | Look-back window in seconds | Defaults to 1800, only the last 30 minutes of errors |
| ids | Strategy input | List of strategy IDs | Limits the scope of monitored strategies, comma-separated |
| status | Strategy input | Execution status code | Fixed value 3, meaning "error" |
| strategy_name | Response field | Name of the failed strategy | Written into the DingTalk message body |
| strategy_id | Response field | ID of the failed strategy | Used for triage and log lookup |
| lessee.name | Response field | Tenant/environment identifier | Distinguishes multi-tenant private deployments |
| number | Response field | Business document number | Lets the business side see which document broke |
| response_at | Response field | Time of the error | Written into the message timestamp for sorting |
| problem | Response field | Exception description | The core content of the DingTalk message |
The source side uses Qeasy's StrategyErrorDetail WebAPI (POST, QUERY); the target side uses DingTalkRobotDetail (POST, EXECUTE). Both are registered under the Qeasy integration platform and configured as "strategy" objects inside it.
How to Configure It in Qeasy
Inside the Qeasy data-integration platform, a common pattern is to split this strategy into a "source strategy" and a "target strategy." Typical configuration points are as follows.
Source side: the error-query strategy
- Pick
StrategyErrorDetailas the API, POST as the method, QUERY as the effect. - Default
recentSecondsto 1800, i.e., only the last half hour, so a backlog does not flood the chat group. - Fill in concrete strategy IDs for
idson first deployment; once stable, you can extend the list or leave it blank. - Fix
statusto3(error). If you also want to see "unreviewed" runs, change it to a multi-value form like3,4. - Keep
autoFillResponse = trueso the platform fills the response shape automatically and manual modelling stays minimal.
Target side: the DingTalk robot push strategy
- Pick
DingTalkRobotDetailas the API, POST as the method, EXECUTE as the effect. access_tokencomes from the DingTalk group custom robot's webhook. Always use a token issued to your own group; do not reuse a screenshot pasted by someone else.- Compose the message body using variable templates.
{{strategy_name}},{{number}},{{response_at}}, and{{problem}}must be kept — these are what an operator needs to triage from the chat. - It is recommended to also include the tenant identifier
{{lessee.name}}in the message body. In multi-tenant private deployments, this avoids the "which environment raised this error?" mystery.
Middle layer: association and triggering
- In Qeasy, link the two strategies via
strategy_name/strategy_id. Each error record returned by the source becomes one execution input on the target. - Code and tenant mappings are best managed centrally in Qeasy's mapping tables rather than scattered across per-strategy scripts. When more strategies are added later, one place to change keeps everything tidy.
Implementation Steps
We recommend rolling it out in three phases: "minimum-viable alert, then expand coverage, then stabilize."
Phase 1: minimum-viable alerting
- Deploy the source error-query strategy. Set
recentSecondsto 1800, keepstatusas3, and leaveidsempty or limited to one or two critical sync strategies. - Point the target DingTalk robot push at a test group first, and confirm that the variables in the message body render correctly.
- Do not route anything to a production group yet. The goal here is to validate the end-to-end path.
Phase 2: expand coverage to all strategies
- Add more strategy IDs to
idson the source side, or keep it empty and let the platform aggregate by tenant. - Switch the target robot to the formal ops group or business-owner group.
- One easy-to-miss detail: stagger the crontab on both sides. The source side is best kept at
*/30 7-23 * * *(every 30 minutes during working hours). The target DingTalk push is best kept at*/5 * * * *(every 5 minutes, so the alert is timely; since real execution only happens when the source has data, it does not spam the group).
Phase 3: steady-state and noise reduction
- After running for a week or two, use group feedback to filter out known-recoverable transient errors from the alert, or aggregate them in scripts.
- Adopt an incremental-and-full dual track: critical sync strategies use the strict alert path, while edge strategies get aggregated into a daily report so that a midnight jitter does not wake people up.
Lessons Learned
- Look-back window too large on day one — group flooded. On the first deployment,
recentSecondswas set to 86400, which pushed days of backlogged errors to the ops group in one shot. The safe approach is small first, larger later: start at 1800, then expand only after confirming it is stable. - DingTalk access_token reused from someone else's screenshot. A typical mistake is to copy a webhook that someone else pasted in the group; the messages then go to the wrong group and ops sees nothing. Every token must be re-copied from your own robot and kept carefully.
- No tenant identifier in multi-tenant environments. Private deployments often host multiple business lines on a single platform. If the message body does not carry
{{lessee.name}}, no one can tell which line is alerting, and triage time doubles. - Source and target schedules are not aligned. If the source polls every 30 minutes and the target pushes every minute, the target will keep spinning even when the source has nothing new. The safe approach is source frequency ≤ target frequency, plus a skip-on-empty-result guard on the target.
- Alerts without document number or timestamp. A DingTalk message that just says "strategy error" still forces the operator back into the platform to read logs. Putting the
{{number}},{{response_at}}, and{{problem}}trio into the message body means the conversation can happen directly in the chat.
When This Fits and When It Does Not
Fits: multi-strategy private-deployment integration environments where exceptions must reach ops or business owners in near real time; cases where you want to fill the observability gap quickly without modifying the business systems themselves. Does not fit: environments with only one or two strategies where the team already reads logs continuously; environments with strict message-compliance or audit rules where any external push must go through an internal approval system — in those cases the compliance boundary of DingTalk-based push must be evaluated first.