Qeasy Cloud
Get Started

Source Platform Special Operations: Flattening Multi-dimensional Data and Queueing Mechanism

· 卢剑航· Product Docs· 3 views· 4 min read

Feature Overview

When designing the Qeasy DataHub data integration platform, we observed that source platform data structures are often highly complex and multi-layered. For instance, an order detail table may contain master table fields along with nested line items (sub-forms). Such two-dimensional or even multi-dimensional data shapes impose a heavy parsing burden during cross-system transmission.

To keep heterogeneous data structurally clear and field-aligned before being written to the target database, we introduced a set of "Special Operations" configuration items in the source platform extension module. The core capabilities currently supported fall into two categories: Flatten Multi-dimensional Data and Queueing. The former addresses data shape transformation, while the latter addresses data write ordering. Both configurations are attached to the source-side node of an integration strategy as JSON snippets, performing pre-processing on the raw payload collected, so that the downstream write logic remains pure.

From an architectural perspective, the Special Operations module sits between the data collector and the field mapping engine, acting as a lightweight "data pre-processing middleware." The original intent behind this design is to decouple cross-cutting concerns—such as multi-dimensional structure handling and write-order control—from the business mapping logic, enabling engineers to complete complex data shaping in a declarative way without writing glue scripts.

Usage Scenarios

Scenario 1: Flattening Multi-dimensional Form Data

A typical use case is when a source platform returns nested JSON via webhooks or polling APIs, such as {"order_id": "X", "details_list": [{"sku": "A", "qty": 1}, {"sku": "B", "qty": 2}]}. If the target system only accepts one-dimensional row data, the details_list array needs to be "flattened" into individual records, with the outer master table fields (e.g., order_id) pushed down into each row. This scenario is extremely common in ERP order synchronization, MES work order flow, and form system integration.

Scenario 2: Source-side Data Write Order Control

The queueing mechanism targets another class of problems: when a source platform pushes data with high concurrency and the business requires downstream consumption to be strictly executed in arrival order or sorted by a particular time field (e.g., state machine transitions, serial number increment), we need to introduce a queue buffer at the source side to prevent data consistency issues caused by out-of-order writes.

Configuration Instructions

How to Configure Flattening Multi-dimensional Data

In the source platform Special Operations configuration panel, after enabling the "Flatten Multi-dimensional Data" switch, you need to declare the nested field paths to be flattened within a JSON configuration block. The configuration structure is illustrated as follows:

json
{
  "beatFlat": [
    "details_list"
  ]
}

The meaning of the above configuration is: split the array field named details_list in the original payload, with each array element becoming an independent output record, while preserving all outer fields as common columns. We support declaring multiple flattening fields at the same time, for example "beatFlat": ["details_list", "attachment_list"]. The system will expand them in the declared order, generating multi-row data in the form of a Cartesian product.

How to Configure Queueing

The queueing module is typically provided as an independent strategy switch. Users can enable queue buffering in the "Advanced Options" of the source-side strategy, and configure the queue capacity, maximum wait time, and sort key (optional). When the queue capacity reaches the threshold, the platform triggers a backpressure mechanism, pausing upstream pulling to ensure no data is lost.

Configuration Activation Flow

  1. Enter the "Integration Strategy" editing page in the Qeasy console;
  2. Select the source platform node and expand the "Special Operations" collapsible panel;
  3. Enable the corresponding feature switch and fill in the configuration according to the JSON Schema;
  4. After saving, use the "Dry Run" feature to verify whether the flattened data shape meets expectations;
  5. After confirmation, publish the strategy and enter production scheduling.

Precautions

  1. Field paths must be precise: The field names in the flattening configuration are JSONPath expressions relative to the payload root node. Incorrect typing or wrong hierarchy will cause the entire batch of data to be discarded. It is recommended to use the preview panel during the dry run stage to verify the number of records after expansion.
  2. Cartesian product risk in multi-dimensional flattening: When flattening multiple array fields at the same time, the output row count is the product of each array's length, which may cause "data explosion." We recommend setting an upper limit in advance for fields with uncontrollable length, or splitting them into different strategies for processing.
  3. Resource consumption of the queueing mechanism: After enabling queueing, a buffer is maintained in memory. Long-term backlog may affect node performance. Please configure the queue capacity reasonably according to actual throughput capability, and configure alert thresholds.
  4. Coupling with field mapping: The field names after flattening will continue to use the keys from the original payload. If the target-side field naming conventions are inconsistent, conversion needs to be completed in the subsequent field mapping step; the two should not be confused.
  5. Version compatibility: Source platform Special Operations is an extended configuration item. When upgrading the Qeasy platform version, please pay attention to the release notes to ensure that the configuration syntax used matches the current version, avoiding historical strategies becoming invalid due to Schema evolution.

By reasonably applying the two source platform special operations—Flatten Multi-dimensional Data and Queueing—engineers can enable DataHub to maintain concise mapping logic and stable execution order even when facing complex source-side structures, further unleashing the engineering efficiency of the Qeasy DataHub data integration platform in heterogeneous data synchronization scenarios.

Original content. Please credit the source when reposting: https://www.qeasy.cloud/insights/product-docs/doc-n523fceaa

Comments