Configure a Data Sync
The Data Sync feature keeps datasets current with their data sources. You can run synchronization manually or on a schedule. Cascade synchronization is also supported when datasets are based on other datasets.
Configure a Data Sync

-
Navigate to Data > Design > Advanced Settings > Data Sync.
The dataset must have been loaded at least once before setting up a sync schedule.
-
Select the Data Sync toggle to enable the feature.
-
Configure Data Source Settings:
-
Off — No synchronization for this data source. Use for static data, like geographical or employee data.
-
Full Reload — Reloads all data. Use when the source lacks unique identifiers or a timestamp column, or to reflect deletions in the dataset.
-
Append and Update — Updates only new or modified records since the last load. While efficient, this method requires specific conditions:
- Unique identifier. Open the Design tab, select Columns, and toggle Unique ID for the column.
- Timestamp column indicating when records were added or updated. Only records with timestamps between the last load and "now minus one minute" are processed.
This method does not delete records that are removed from the source.
-
Custom - Enables you to customize the Append and Update query. Use this feature to optimize the default Append and Update query for better performance.
Note: If performance degrades over time, consider optimizing the join lake.
-
-
Select Sync Now for a manual sync.
-
Under Schedule, configure sync scheduling options.
-
Select Apply Changes to save the configuration.
-
Repeat these steps for each data source in the dataset.
- Data remains accessible during synchronization as the process runs in the background.
- Refreshed data is only available after the sync finishes.
- Multiple synchronizations for the same dataset cannot run simultaneously.
- You can also use Qrvey's Automated Workflows or APIs to synchronize data.
Schedule a Data Sync
You can schedule syncs using the following methods:
- Basic: Set schedules using the dropdowns.
- CRON: Use CRON expressions for advanced scheduling. Test the expression to ensure proper syntax. For more information, see AWS CRON Documentation.
Tips:
- Schedule syncs during low-traffic hours to reduce resource strain.
- Stagger sync schedules to avoid conflicts.
Configure a Cascade Synchronization
Cascade synchronization triggers a sync automatically when underlying datasets are updated. No schedule is required.
- Under Schedule, select When data sources are updated.
- Ensure the dataset contains at least one other dataset as a data source.
Run a Manual Sync
To run a manual synchronization:
- Open the dataset and go to Data Sync.
- Toggle Data Sync to activate it.
- Configure the Sync Type under Data Source Settings.
- Select Sync Now. The sync runs immediately, and scheduled syncs continue as configured.
Determine Synchronization Logic
-
Identify the refresh schedule. If no schedule is required, consider manual updates or automation triggers.
-
Determine the best synchronization type for each data source.
- For static data sources (no refresh needed), keep Sync Type set to Off.
- If the data source has a unique identifier and timestamp:
- Use Append and Update when updates need to be reflected.
- Use Custom when the Append and Update query needs to be optimized for better performance.
- Use Full Reload to replace the entire dataset or when the data source is missing an identifier or timestamp.
-
Adjust the query start time (if needed).
Usually, defaults are appropriate, but adjust for historical data added before the last load. To ensure inclusion, set the Query Start Time to the earliest record's timestamp.
-
Select the synchronization frequency.
- Keep syncs as infrequent as possible to reduce resource consumption.
- Schedule during off-hours or low-traffic periods.
-
Toggle Data Sync on to start the synchronization schedule, or select Sync Now to run the synchronization immediately.
-
Select Apply Changes to persist your configuration.
You can monitor the synchronization performance and adjust the schedule or logic as needed.
Custom Queries in Data Syncs
Qrvey supports the use of custom queries in a data sync for retrieving recent records. When a dataset sync is using Append and Update, Qrvey normally takes your original query and wraps it with additional logic to retrieve only the most recently changed records. This works for simple queries, but the additional wrapping can slow down syncs for more complex queries (joins, unions, subqueries).
Use Case
Consider adding a custom sync to a data source to address the following needs:
- Performance issues caused by wrapping the original complex query (for example, unions across multiple tables) in a timestamp filter.
- You want full control over how “most recent records” are determined, rather than relying on Qrvey's automatic wrapping.
Custom sync queries are available for data sources that support MongoDB, DynamoDB, and SQL/custom queries when the dataset has an ID column defined.
Each data source in a dataset can have its own sync configuration. For example, DataSourceA might use Full Reload, while DataSourceB uses Append and Update, and DataSourceC uses Custom.
Add a Custom Sync Query
-
From the data source, set the sync type to Custom.
Qrvey wraps your original query with start and end date filters, the same logic used by the default Append and Update sync.
-
Edit this query as needed.
Unlike a standard query, a custom sync query does not require timestamp columns. You define how the timeframe is applied, and you can cast or transform columns differently if needed.
-
To insert a time boundary, type
{{in the query editor. Qrvey displays the available tokens for the start and end of the sync window (for example, the last successful sync time and the current time).When testing the query, you can supply sample values for these tokens. When a real sync runs, Qrvey substitutes the actual sync start and end times automatically.
-
If needed, enable the long-running query option for the initial sync query. An initial load often takes longer than a day-to-day sync.
-
Test the query before saving. Qrvey checks the columns returned against your original query and dataset and prompts you if it finds differences such as new or removed columns, changed data types, or different tables or views. You can continue if the difference is expected.
Edit a Custom Sync Query
After creating a custom sync query, you can edit it in the following areas:
- Data Sync tab
- Data source pill (edit query option)
Changes to Original Query
If you edit the original query for a data source that already has a custom sync query defined, Qrvey displays a banner prompting you about a possible mismatch. Review and update the custom sync query if needed.
Data Fragmentation and Join Lake Optimization
When you set up an incremental data sync, new data and updates are stored in new files while older files are retained. Over time, this leads to data fragmentation (many small files containing pieces of your dataset). This fragmentation can slow data joins and the sync process response, especially as the dataset grows. The implementation of join lake optimization is specific to optimizing S3 Lakes for the Joins feature in Qrvey.
To address this, periodically run the Join Lake Optimization process. This optimization consolidates fragmented data into fewer, larger files, improving performance for data joins and syncs.
For best results, run Join Lake Optimization after a significant number of incremental syncs have occurred and your dataset has grown substantially. For most customers, this means running the process periodically (such as monthly), or after major data updates—to maintain optimal performance. If you notice slower data joins or syncs, consider running the optimization more frequently.
For more information, see API documentation.
Run a Join Lake Optimization
The Join Lake Optimization requires that at least one data source in your dataset uses the Append and Update sync setting.
You can trigger Join Lake Optimization for a specific dataset by calling the API.
API Endpoint:
POST https://{composer_url}/devapi/v5/dataset/jlo
Content-Type: application/json
x-api-key: <your_api_key>
{
"datasetId": "<your_dataset_id>"
}
cURL Example:
curl --location 'https://{composer_url}/devapi/v5/dataset/jlo' \
--header 'Content-Type: application/json' \
--header 'x-api-key: <your_api_key>' \
--data '{
"datasetId":"<your_dataset_id>"
}'
Response Example:
{
"ok": true,
"message": "Join Lake Optimization Process Has Started.",
"datasetId": "<your_dataset_id>",
"datasetName": "<dataset>"
}
After triggering the process, check the Activity Log for your dataset to monitor the start and completion of the optimization.