# Data Preparation

**URL:** https://community.sparkflows.ai/c/data-preparation/10.md

[Latest](https://community.sparkflows.ai/latest.md) · [Categories](https://community.sparkflows.ai/categories.md) · [Tags](https://community.sparkflows.ai/tags.md)

---

## [About the Data Preparation category](https://community.sparkflows.ai/t/about-the-data-preparation-category/16)

<div class="topic-metadata">

**Author:** [@admin](https://community.sparkflows.ai/u/admin)\
**Replies:** 0

</div>

---

## [How to Rename Dynamic Year-Based Columns and Pass Column Lists as Parameters in Sparkflows?](https://community.sparkflows.ai/t/how-to-rename-dynamic-year-based-columns-and-pass-column-lists-as-parameters-in-sparkflows/310)

<div class="topic-metadata">

**Author:** [@admin](https://community.sparkflows.ai/u/admin)\
**Replies:** 0\
**Last updated:** [March 2, 2026, 6:41am UTC](https://community.sparkflows.ai/t/how-to-rename-dynamic-year-based-columns-and-pass-column-lists-as-parameters-in-sparkflows/310 "2026-03-02T06:41:46Z")

</div>

Question The column name changes every year. For example: Sum of Q1'25, Sum of Q2'25, Sum of Q3'25, Sum of Q4'25, Sum of FY'25 For year 2026, it becomes: Sum of Q1'26 How do we read and rename the columns to a fixed …

---

## [How to write conditions in the GE Decision node?](https://community.sparkflows.ai/t/how-to-write-conditions-in-the-ge-decision-node/166)

<div class="topic-metadata">

**Author:** [@Tarika](https://community.sparkflows.ai/u/Tarika)\
**Replies:** 0\
**Last updated:** [December 18, 2025, 11:39am UTC](https://community.sparkflows.ai/t/how-to-write-conditions-in-the-ge-decision-node/166 "2025-12-18T11:39:09Z")

</div>

To write conditions in the GE Decision node, follow these steps: The GE Decision node requires an input DataFrame from a CSV file created from GE Results. This is typically done using the Create CSV from the GE Results …

---

## [I want to get the row count at a given stage. How to achieve this in Sparkflows?](https://community.sparkflows.ai/t/i-want-to-get-the-row-count-at-a-given-stage-how-to-achieve-this-in-sparkflows/165)

<div class="topic-metadata">

**Author:** [@Tarika](https://community.sparkflows.ai/u/Tarika)\
**Replies:** 0\
**Last updated:** [December 18, 2025, 11:24am UTC](https://community.sparkflows.ai/t/i-want-to-get-the-row-count-at-a-given-stage-how-to-achieve-this-in-sparkflows/165 "2025-12-18T11:24:59Z")

</div>

In Sparkflows, we can use the Count processor to get row count at any stage. It can be attached to a processor where count is to be derived and it would print the row count. To use the ‘Count’ Processor, do the followin…

---

## [I have an employee department dataset containing salary information. I want to identify the minimum and maximum salary for each department and location. How to achieve this in Sparkflows?](https://community.sparkflows.ai/t/i-have-an-employee-department-dataset-containing-salary-information-i-want-to-identify-the-minimum-and-maximum-salary-for-each-department-and-location-how-to-achieve-this-in-sparkflows/164)

<div class="topic-metadata">

**Author:** [@Tarika](https://community.sparkflows.ai/u/Tarika)\
**Replies:** 0\
**Last updated:** [December 18, 2025, 11:07am UTC](https://community.sparkflows.ai/t/i-have-an-employee-department-dataset-containing-salary-information-i-want-to-identify-the-minimum-and-maximum-salary-for-each-department-and-location-how-to-achieve-this-in-sparkflows/164 "2025-12-18T11:07:36Z")

</div>

In Sparkflows, we can use the Multi Windows Analytics processor to compute minimum and maximum values. First it would create a partition by department and location. Then it would compute minimum and maximum salary values …

---

## [I have an employee department dataset containing salary information. I want to get a salary based ranking within each department and location. How to achieve this in Sparkflows?](https://community.sparkflows.ai/t/i-have-an-employee-department-dataset-containing-salary-information-i-want-to-get-a-salary-based-ranking-within-each-department-and-location-how-to-achieve-this-in-sparkflows/163)

<div class="topic-metadata">

**Author:** [@Tarika](https://community.sparkflows.ai/u/Tarika)\
**Replies:** 0\
**Last updated:** [December 18, 2025, 11:03am UTC](https://community.sparkflows.ai/t/i-have-an-employee-department-dataset-containing-salary-information-i-want-to-get-a-salary-based-ranking-within-each-department-and-location-how-to-achieve-this-in-sparkflows/163 "2025-12-18T11:03:42Z")

</div>

In Sparkflows, we can use the Multi Windows Ranking processor to get ranking within a partition. First it would create a partition by department and location. Then it would rank based on salary. To use the ‘Multi Window…

---

## [In my dataset I have salaries for employees containing various decimal values. I want to round it off to 2 decimal values. How to achieve this in Sparkflows?](https://community.sparkflows.ai/t/in-my-dataset-i-have-salaries-for-employees-containing-various-decimal-values-i-want-to-round-it-off-to-2-decimal-values-how-to-achieve-this-in-sparkflows/162)

<div class="topic-metadata">

**Author:** [@Tarika](https://community.sparkflows.ai/u/Tarika)\
**Replies:** 0\
**Last updated:** [December 18, 2025, 10:50am UTC](https://community.sparkflows.ai/t/in-my-dataset-i-have-salaries-for-employees-containing-various-decimal-values-i-want-to-round-it-off-to-2-decimal-values-how-to-achieve-this-in-sparkflows/162 "2025-12-18T10:50:17Z")

</div>

In Sparkflows, you can use the Round Value processor to round off values to desired decimal places. To use the ‘Round Value’ Processor, do the following: Select columns to be rounded off in ‘Input Column’. It can be …

---

## [How to split the string value in a column into multiple columns in Sparkflows?](https://community.sparkflows.ai/t/how-to-split-the-string-value-in-a-column-into-multiple-columns-in-sparkflows/161)

<div class="topic-metadata">

**Author:** [@Tarika](https://community.sparkflows.ai/u/Tarika)\
**Replies:** 0\
**Last updated:** [December 18, 2025, 10:35am UTC](https://community.sparkflows.ai/t/how-to-split-the-string-value-in-a-column-into-multiple-columns-in-sparkflows/161 "2025-12-18T10:35:31Z")

</div>

Sparkflows provides the node called Field Splitter to create multiple columns from a single column provided by a separator. To know more, refer here: Parse Functions — Sparkflows 3.0 documentation

---

## [How to split the input data into two outputs depending on the condition?](https://community.sparkflows.ai/t/how-to-split-the-input-data-into-two-outputs-depending-on-the-condition/160)

<div class="topic-metadata">

**Author:** [@Tarika](https://community.sparkflows.ai/u/Tarika)\
**Replies:** 0\
**Last updated:** [December 18, 2025, 10:32am UTC](https://community.sparkflows.ai/t/how-to-split-the-input-data-into-two-outputs-depending-on-the-condition/160 "2025-12-18T10:32:22Z")

</div>

Sparkflows provide the node called Split By Expression to split the input DataFrame into two DataFrames. To know more, read this documentation: Split Dataset By Expression — Sparkflows 3.0 documentation

---

## [Given a retail dataset, the goal is to validate the schema, address fields, and overall structure to ensure data accuracy and quality for reliable analysis and decision-making. How can this be achieved?](https://community.sparkflows.ai/t/given-a-retail-dataset-the-goal-is-to-validate-the-schema-address-fields-and-overall-structure-to-ensure-data-accuracy-and-quality-for-reliable-analysis-and-decision-making-how-can-this-be-achieved/159)

<div class="topic-metadata">

**Author:** [@Tarika](https://community.sparkflows.ai/u/Tarika)\
**Replies:** 0\
**Last updated:** [December 18, 2025, 10:23am UTC](https://community.sparkflows.ai/t/given-a-retail-dataset-the-goal-is-to-validate-the-schema-address-fields-and-overall-structure-to-ensure-data-accuracy-and-quality-for-reliable-analysis-and-decision-making-how-can-this-be-achieved/159 "2025-12-18T10:23:07Z")

</div>

Sparkflows provides a diverse range of nodes that cater to the above-mentioned requirements as these nodes assist in ensuring data quality and integrity. Some of them are listed below : Node Schema Validation Valid…

---

## [How can I explore data with the help of Sparkflows?](https://community.sparkflows.ai/t/how-can-i-explore-data-with-the-help-of-sparkflows/158)

<div class="topic-metadata">

**Author:** [@Tarika](https://community.sparkflows.ai/u/Tarika)\
**Replies:** 0\
**Last updated:** [December 18, 2025, 10:01am UTC](https://community.sparkflows.ai/t/how-can-i-explore-data-with-the-help-of-sparkflows/158 "2025-12-18T10:01:51Z")

</div>

Sparkflows offers a range of nodes for data profiling and exploratory analysis, allowing users to examine data profiles and perform comprehensive exploration tasks. Those are listed below: Summary Statistics Columns Ca…

---

## [I possess EHR data which contains nested JSON format. I would like to extract all the fields and convert them into a CSV file. How can I achieve this in Sparkflows?](https://community.sparkflows.ai/t/i-possess-ehr-data-which-contains-nested-json-format-i-would-like-to-extract-all-the-fields-and-convert-them-into-a-csv-file-how-can-i-achieve-this-in-sparkflows/157)

<div class="topic-metadata">

**Author:** [@Tarika](https://community.sparkflows.ai/u/Tarika)\
**Replies:** 0\
**Last updated:** [December 18, 2025, 9:50am UTC](https://community.sparkflows.ai/t/i-possess-ehr-data-which-contains-nested-json-format-i-would-like-to-extract-all-the-fields-and-convert-them-into-a-csv-file-how-can-i-achieve-this-in-sparkflows/157 "2025-12-18T09:50:54Z")

</div>

You can use the Flatten and Explode nodes. The “Flatten” node allows to convert complex nested structures into a more straightforward columnar format. On the other hand, the “Explode” node is useful for breaking down arr…

---

## [I want to use only a smaller set of data for my analysis. How to achieve this in Sparkflows?](https://community.sparkflows.ai/t/i-want-to-use-only-a-smaller-set-of-data-for-my-analysis-how-to-achieve-this-in-sparkflows/156)

<div class="topic-metadata">

**Author:** [@Tarika](https://community.sparkflows.ai/u/Tarika)\
**Replies:** 0\
**Last updated:** [December 18, 2025, 9:45am UTC](https://community.sparkflows.ai/t/i-want-to-use-only-a-smaller-set-of-data-for-my-analysis-how-to-achieve-this-in-sparkflows/156 "2025-12-18T09:45:30Z")

</div>

In Sparkflows, we can use the Sample processor to extract a sample of incoming datasets. The number of rows in the sample would be a percentage of the incoming dataset. Sample can be used for ML Training or Analysis purp…

---

## [I want to sort incoming dataset. How to achieve this in Sparkflows?](https://community.sparkflows.ai/t/i-want-to-sort-incoming-dataset-how-to-achieve-this-in-sparkflows/155)

<div class="topic-metadata">

**Author:** [@Tarika](https://community.sparkflows.ai/u/Tarika)\
**Replies:** 0\
**Last updated:** [December 18, 2025, 9:32am UTC](https://community.sparkflows.ai/t/i-want-to-sort-incoming-dataset-how-to-achieve-this-in-sparkflows/155 "2025-12-18T09:32:35Z")

</div>

In Sparkflows, we can use the Sort By processor to sort incoming dataset. Dataset can be sorted based on one or multiple columns. To use the ‘Sort By’ Processor: Select a column and sorting order to sort the incoming…

---

## [How can I transpose a dataset in Sparkflows?](https://community.sparkflows.ai/t/how-can-i-transpose-a-dataset-in-sparkflows/145)

<div class="topic-metadata">

**Author:** [@Tarika](https://community.sparkflows.ai/u/Tarika)\
**Replies:** 0\
**Last updated:** [December 17, 2025, 11:46am UTC](https://community.sparkflows.ai/t/how-can-i-transpose-a-dataset-in-sparkflows/145 "2025-12-17T11:46:54Z")

</div>

In Sparkflows, we can use the Transpose processor to transpose a dataset. To use the ‘Transpose’ Processor: Select a column to be used to transpose incoming dataset in ‘Transpose By Column Name’. Output would be a…

---

## [I want to sort incoming columns in a particular order](https://community.sparkflows.ai/t/i-want-to-sort-incoming-columns-in-a-particular-order/144)

<div class="topic-metadata">

**Author:** [@Tarika](https://community.sparkflows.ai/u/Tarika)\
**Replies:** 0\
**Last updated:** [December 17, 2025, 11:40am UTC](https://community.sparkflows.ai/t/i-want-to-sort-incoming-columns-in-a-particular-order/144 "2025-12-17T11:40:43Z")

</div>

In Sparkflows, you can use the Sort Columns processor to sort incoming columns in any order. You can use the ‘Sort Columns’ processor as below: Columns can be sorted in Ascending or Descending order of column names. C…

---

## [I want to perform state-wise analysis on a dataset related to population](https://community.sparkflows.ai/t/i-want-to-perform-state-wise-analysis-on-a-dataset-related-to-population/143)

<div class="topic-metadata">

**Author:** [@Tarika](https://community.sparkflows.ai/u/Tarika)\
**Replies:** 0\
**Last updated:** [December 17, 2025, 11:33am UTC](https://community.sparkflows.ai/t/i-want-to-perform-state-wise-analysis-on-a-dataset-related-to-population/143 "2025-12-17T11:33:12Z")

</div>

In Sparkflows, we can use the Windows Aggregation processor to perform state wise data analysis. Data can be partitioned by State and various computations can be performed such as avg, max, min, range, std dev, and so on…

---

## [How can I perform Windows Analytics on a dataset in Sparkflows?](https://community.sparkflows.ai/t/how-can-i-perform-windows-analytics-on-a-dataset-in-sparkflows/142)

<div class="topic-metadata">

**Author:** [@Tarika](https://community.sparkflows.ai/u/Tarika)\
**Replies:** 0\
**Last updated:** [December 17, 2025, 11:28am UTC](https://community.sparkflows.ai/t/how-can-i-perform-windows-analytics-on-a-dataset-in-sparkflows/142 "2025-12-17T11:28:50Z")

</div>

In Sparkflows, we can use the ‘Windows Analytics’ processor to perform Windows Analytics on a dataset. It facilitates partitioning the dataset based on a selected column and applies windows functions such as first\_value,…

---

## [How do I split my data into unique and duplicate records in Sparkflows?](https://community.sparkflows.ai/t/how-do-i-split-my-data-into-unique-and-duplicate-records-in-sparkflows/141)

<div class="topic-metadata">

**Author:** [@Tarika](https://community.sparkflows.ai/u/Tarika)\
**Replies:** 0\
**Last updated:** [December 17, 2025, 10:52am UTC](https://community.sparkflows.ai/t/how-do-i-split-my-data-into-unique-and-duplicate-records-in-sparkflows/141 "2025-12-17T10:52:50Z")

</div>

In Sparkflows, there is a Find Duplicate node that can perform the exact operation you’re looking for. You can specify the column(s) based on which you want to determine uniqueness. Simply add this node to your input dat…

---

## [Is there a way to replace a specific value in a column using Sparkflows?](https://community.sparkflows.ai/t/is-there-a-way-to-replace-a-specific-value-in-a-column-using-sparkflows/140)

<div class="topic-metadata">

**Author:** [@Tarika](https://community.sparkflows.ai/u/Tarika)\
**Replies:** 0\
**Last updated:** [December 17, 2025, 9:42am UTC](https://community.sparkflows.ai/t/is-there-a-way-to-replace-a-specific-value-in-a-column-using-sparkflows/140 "2025-12-17T09:42:34Z")

</div>

To replace a specific value in a column using Sparkflows, you can utilize the Impute Advanced node, which allows you to configure value replacements or imputations. To replace a specific value with a constant, you simpl…

---

## [How can I filter good and bad records after performing data quality checks using Great Expectation nodes?](https://community.sparkflows.ai/t/how-can-i-filter-good-and-bad-records-after-performing-data-quality-checks-using-great-expectation-nodes/139)

<div class="topic-metadata">

**Author:** [@Tarika](https://community.sparkflows.ai/u/Tarika)\
**Replies:** 0\
**Last updated:** [December 17, 2025, 8:36am UTC](https://community.sparkflows.ai/t/how-can-i-filter-good-and-bad-records-after-performing-data-quality-checks-using-great-expectation-nodes/139 "2025-12-17T08:36:29Z")

</div>

To separate good and bad records, you can utilize the "Split Into Good Bad Records’’ node. After adding this node to your workflow following any Great Expectation node, ensure that the input DataFrame is the original Dat…

---

## [How to Normalize data using Sparkflows?](https://community.sparkflows.ai/t/how-to-normalize-data-using-sparkflows/130)

<div class="topic-metadata">

**Author:** [@Tarika](https://community.sparkflows.ai/u/Tarika)\
**Replies:** 0\
**Last updated:** [December 16, 2025, 1:05pm UTC](https://community.sparkflows.ai/t/how-to-normalize-data-using-sparkflows/130 "2025-12-16T13:05:47Z")

</div>

In Sparkflows, you can use the “Normalizer” node to normalize the data. Let us try to understand the process with the help of a simple workflow. Below is the image of a simple workflow using the Normalizer node. Bef…

---

## [I want to get word count from the reviews before they are fed to the database. How can I achieve this in sparkflows?](https://community.sparkflows.ai/t/i-want-to-get-word-count-from-the-reviews-before-they-are-fed-to-the-database-how-can-i-achieve-this-in-sparkflows/129)

<div class="topic-metadata">

**Author:** [@Tarika](https://community.sparkflows.ai/u/Tarika)\
**Replies:** 0\
**Last updated:** [December 16, 2025, 12:44pm UTC](https://community.sparkflows.ai/t/i-want-to-get-word-count-from-the-reviews-before-they-are-fed-to-the-database-how-can-i-achieve-this-in-sparkflows/129 "2025-12-16T12:44:36Z")

</div>

In Sparkflows, we can use the ‘Word Count’ processor to get word count of selected columns. To use the ‘Word Count’ Processor: Select a set of columns for which word count is to be computed in the ‘Input Columns’ fiel…

---

## [I want to analyze log data to get meaningful insight. How can I achieve this in sparkflows?](https://community.sparkflows.ai/t/i-want-to-analyze-log-data-to-get-meaningful-insight-how-can-i-achieve-this-in-sparkflows/128)

<div class="topic-metadata">

**Author:** [@Tarika](https://community.sparkflows.ai/u/Tarika)\
**Replies:** 0\
**Last updated:** [December 16, 2025, 12:38pm UTC](https://community.sparkflows.ai/t/i-want-to-analyze-log-data-to-get-meaningful-insight-how-can-i-achieve-this-in-sparkflows/128 "2025-12-16T12:38:50Z")

</div>

In Sparkflows, we can use the “Apache Logs” processor to read and process log data. It reads a log file and loads it as a DataFrame. Thereafter DataFrame can be used for further analysis. To use the “Apache Logs” Proces…

---

## [I have one year's worth of sales data, and I'd like to calculate the cumulative sales total, enabling me to analyze the total sales made up to a specific date](https://community.sparkflows.ai/t/i-have-one-years-worth-of-sales-data-and-id-like-to-calculate-the-cumulative-sales-total-enabling-me-to-analyze-the-total-sales-made-up-to-a-specific-date/127)

<div class="topic-metadata">

**Author:** [@Tarika](https://community.sparkflows.ai/u/Tarika)\
**Replies:** 0\
**Last updated:** [December 16, 2025, 12:31pm UTC](https://community.sparkflows.ai/t/i-have-one-years-worth-of-sales-data-and-id-like-to-calculate-the-cumulative-sales-total-enabling-me-to-analyze-the-total-sales-made-up-to-a-specific-date/127 "2025-12-16T12:31:37Z")

</div>

You can fulfill the aforementioned requirement by utilizing a window aggregate function through Window Aggregation Node, enabling you to compute the running total of sales. Moreover, you can also calculate the cumulative…

---

## [How does Apache Spark’s distributed execution affect the number of output files, and what method ensures saving the result in only one file?](https://community.sparkflows.ai/t/how-does-apache-spark-s-distributed-execution-affect-the-number-of-output-files-and-what-method-ensures-saving-the-result-in-only-one-file/89)

<div class="topic-metadata">

**Author:** [@Tarika](https://community.sparkflows.ai/u/Tarika)\
**Replies:** 0\
**Last updated:** [December 12, 2025, 2:13pm UTC](https://community.sparkflows.ai/t/how-does-apache-spark-s-distributed-execution-affect-the-number-of-output-files-and-what-method-ensures-saving-the-result-in-only-one-file/89 "2025-12-12T14:13:50Z")

</div>

Apache Spark runs distributed. As a result the data is partitioned across multiple process/machines. When any of the Save nodes is used to write the output to files, the number of files created is dependent on the numbe…

---

## [Regex used to add \_ in front number column name](https://community.sparkflows.ai/t/regex-used-to-add-in-front-number-column-name/80)

<div class="topic-metadata">

**Author:** [@jayant](https://community.sparkflows.ai/u/jayant)\
**Replies:** 0\
**Last updated:** [December 12, 2025, 9:32am UTC](https://community.sparkflows.ai/t/regex-used-to-add-in-front-number-column-name/80 "2025-12-12T09:32:01Z")

</div>

BulkColumnRename node can be used to add \_ in front number column name.

---

## [How do I aggregate columns in the workflow designer?](https://community.sparkflows.ai/t/how-do-i-aggregate-columns-in-the-workflow-designer/75)

<div class="topic-metadata">

**Author:** [@jayant](https://community.sparkflows.ai/u/jayant)\
**Replies:** 0\
**Last updated:** [December 11, 2025, 7:38pm UTC](https://community.sparkflows.ai/t/how-do-i-aggregate-columns-in-the-workflow-designer/75 "2025-12-11T19:38:19Z")

</div>

Columns can be aggregated using the Aggregate node. You can provide an expression for the aggregation. You can also name the aggregated column in the process.

---

## [I would like to extract currency value from this column. How can I achieve this in Sparkflows? I have a dataset containing an amount column in the format ccy + value (USD 1000000.00)](https://community.sparkflows.ai/t/i-would-like-to-extract-currency-value-from-this-column-how-can-i-achieve-this-in-sparkflows-i-have-a-dataset-containing-an-amount-column-in-the-format-ccy-value-usd-1000000-00/68)

<div class="topic-metadata">

**Author:** [@Ragita](https://community.sparkflows.ai/u/Ragita)\
**Replies:** 0\
**Last updated:** [December 11, 2025, 9:00am UTC](https://community.sparkflows.ai/t/i-would-like-to-extract-currency-value-from-this-column-how-can-i-achieve-this-in-sparkflows-i-have-a-dataset-containing-an-amount-column-in-the-format-ccy-value-usd-1000000-00/68 "2025-12-11T09:00:09Z")

</div>

In Sparkflows, we can use the “Multi Regex Extractor” processor to achieve this. Processor needs to be configured as below and it would extract the currency part. To use the “Multi Regex Extractor” Processor: Select …

---

## [I have a dataset having sales information of multiple stores from various locations. How can I rank stores by their sales values within a locality using sparkflows?](https://community.sparkflows.ai/t/i-have-a-dataset-having-sales-information-of-multiple-stores-from-various-locations-how-can-i-rank-stores-by-their-sales-values-within-a-locality-using-sparkflows/41)

<div class="topic-metadata">

**Author:** [@Ragita](https://community.sparkflows.ai/u/Ragita)\
**Replies:** 0\
**Last updated:** [December 10, 2025, 6:20am UTC](https://community.sparkflows.ai/t/i-have-a-dataset-having-sales-information-of-multiple-stores-from-various-locations-how-can-i-rank-stores-by-their-sales-values-within-a-locality-using-sparkflows/41 "2025-12-10T06:20:00Z")

</div>

In Sparkflows, we can use the ‘Windows Ranking’ processor to rank stores based on their sales value within a location. Dataset can be partitioned by location and sorted by sales values. Using this configuration we can ge…

[Next page](https://community.sparkflows.ai/c/data-preparation/10.md?page=1)
