# \#faq

**URL:** https://community.sparkflows.ai/tag/faq/3.md

[Latest](https://community.sparkflows.ai/latest.md) · [Categories](https://community.sparkflows.ai/categories.md) · [Tags](https://community.sparkflows.ai/tags.md)

---

## [How to import from another git repository](https://community.sparkflows.ai/t/how-to-import-from-another-git-repository/316)

<div class="topic-metadata">

**Author:** [@Daniel](https://community.sparkflows.ai/u/Daniel)\
**Replies:** 0\
**Last updated:** [March 3, 2026, 5:41am UTC](https://community.sparkflows.ai/t/how-to-import-from-another-git-repository/316 "2026-03-03T05:41:45Z")

</div>

Importing from Another Git Repository 1. What does “Import from other repo” mean? This option allows you to import a project from any external Git repository by providing its repository URL. 2. How do I import a projec…

---

## [Importing from a Pre-configured Global Git Repository](https://community.sparkflows.ai/t/importing-from-a-pre-configured-global-git-repository/315)

<div class="topic-metadata">

**Author:** [@Daniel](https://community.sparkflows.ai/u/Daniel)\
**Replies:** 0\
**Last updated:** [March 2, 2026, 11:55am UTC](https://community.sparkflows.ai/t/importing-from-a-pre-configured-global-git-repository/315 "2026-03-02T11:55:26Z")

</div>

1. What does “Import from global repo” mean? This option allows you to import a project from a Git repository that has already been configured globally by the admin in Sparkflows. 2. How do I import a project from the …

---

## [How to Import a Project from Git in Sparkflows](https://community.sparkflows.ai/t/how-to-import-a-project-from-git-in-sparkflows/314)

<div class="topic-metadata">

**Author:** [@Daniel](https://community.sparkflows.ai/u/Daniel)\
**Replies:** 0\
**Last updated:** [March 2, 2026, 11:54am UTC](https://community.sparkflows.ai/t/how-to-import-a-project-from-git-in-sparkflows/314 "2026-03-02T11:54:16Z")

</div>

1. What is the purpose of the “Import Project” feature in Sparkflows? The Import Project feature in Sparkflows allows users to bring existing projects from a Git repository directly into the platform. This makes it easy …

---

## [How do you configure Git credentials in Sparkflows?](https://community.sparkflows.ai/t/how-do-you-configure-git-credentials-in-sparkflows/309)

<div class="topic-metadata">

**Author:** [@Daniel](https://community.sparkflows.ai/u/Daniel)\
**Replies:** 0\
**Last updated:** [February 27, 2026, 7:37am UTC](https://community.sparkflows.ai/t/how-do-you-configure-git-credentials-in-sparkflows/309 "2026-02-27T07:37:25Z")

</div>

How do you configure Git credentials in Sparkflows? Follow these steps: Click on the Waffle icon (nine squares) at the top-right corner. Select Git Configuration. Enter the Username for the Git account. Enter…

---

## [How do you enable Git in Sparkflows?](https://community.sparkflows.ai/t/how-do-you-enable-git-in-sparkflows/308)

<div class="topic-metadata">

**Author:** [@Daniel](https://community.sparkflows.ai/u/Daniel)\
**Replies:** 0\
**Last updated:** [February 27, 2026, 7:36am UTC](https://community.sparkflows.ai/t/how-do-you-enable-git-in-sparkflows/308 "2026-02-27T07:36:24Z")

</div>

How do you enable Git in Sparkflows? Follow these steps: Login as Admin. Click on Administration from the top menu. Select Configurations. Click on the GIT tab. Set the property git.enabled to true. Pro…

---

## [What is Git Configuration in Sparkflows?](https://community.sparkflows.ai/t/what-is-git-configuration-in-sparkflows/307)

<div class="topic-metadata">

**Author:** [@Daniel](https://community.sparkflows.ai/u/Daniel)\
**Replies:** 0\
**Last updated:** [February 26, 2026, 9:40am UTC](https://community.sparkflows.ai/t/what-is-git-configuration-in-sparkflows/307 "2026-02-26T09:40:35Z")

</div>

1. What is Git Configuration in Sparkflows? Git Configuration in Sparkflows is the process of enabling and setting up Git integration so that artifacts such as projects, workflows, and pipelines can be stored and version…

---

## [How does Git Integration work in Sparkflows and what is the purpose](https://community.sparkflows.ai/t/how-does-git-integration-work-in-sparkflows-and-what-is-the-purpose/304)

<div class="topic-metadata">

**Author:** [@Daniel](https://community.sparkflows.ai/u/Daniel)\
**Replies:** 0\
**Last updated:** [February 26, 2026, 4:13am UTC](https://community.sparkflows.ai/t/how-does-git-integration-work-in-sparkflows-and-what-is-the-purpose/304 "2026-02-26T04:13:18Z")

</div>

1. What is Git Integration in Sparkflows? Git Integration in Sparkflows allows you to persist platform artifacts—such as projects, datasets, workflows, and pipelines—to a Git repository. This enables proper version contr…

---

## [What are the common macros that can be used in the Sparkflows with examples](https://community.sparkflows.ai/t/what-are-the-common-macros-that-can-be-used-in-the-sparkflows-with-examples/303)

<div class="topic-metadata">

**Author:** [@Daniel](https://community.sparkflows.ai/u/Daniel)\
**Replies:** 0\
**Last updated:** [February 25, 2026, 9:19am UTC](https://community.sparkflows.ai/t/what-are-the-common-macros-that-can-be-used-in-the-sparkflows-with-examples/303 "2026-02-25T09:19:05Z")

</div>

{{ ds }} returns the execution date in YYYY-MM-DD format. 2018-01-08 {{ ds\_nodash }} returns the execution date in YYYYMMDD format. 20180108 {{ ts }} returns the execution timestamp in ISO format. 2018-01-01T00:00:00+0…

---

## [What is a macro in Sparkfllows?](https://community.sparkflows.ai/t/what-is-a-macro-in-sparkfllows/302)

<div class="topic-metadata">

**Author:** [@Daniel](https://community.sparkflows.ai/u/Daniel)\
**Replies:** 0\
**Last updated:** [February 25, 2026, 8:56am UTC](https://community.sparkflows.ai/t/what-is-a-macro-in-sparkfllows/302 "2026-02-25T08:56:55Z")

</div>

A macro in Sparkflows is a variable that gets expanded into a string at runtime. Macros enable dynamic template generation, allowing you to insert execution-specific values into SQL queries, file paths, commands, and oth…

---

## [How do I use Copilot to build a workflow from scratch?](https://community.sparkflows.ai/t/how-do-i-use-copilot-to-build-a-workflow-from-scratch/300)

<div class="topic-metadata">

**Author:** [@Aayush](https://community.sparkflows.ai/u/Aayush)\
**Replies:** 0\
**Last updated:** [February 4, 2026, 6:12pm UTC](https://community.sparkflows.ai/t/how-do-i-use-copilot-to-build-a-workflow-from-scratch/300 "2026-02-04T18:12:19Z")

</div>

1. Open the Workflow Designer and click the Copilot button on the top-right toolbar. In the Assistant panel that appears on the right, type your prompt (e.g., “Read Parquet from S3, filter price \> 40000, and save t…

---

## [How do I connect Copilot to a GENAI connection?](https://community.sparkflows.ai/t/how-do-i-connect-copilot-to-a-genai-connection/299)

<div class="topic-metadata">

**Author:** [@Aayush](https://community.sparkflows.ai/u/Aayush)\
**Replies:** 0\
**Last updated:** [February 4, 2026, 5:35pm UTC](https://community.sparkflows.ai/t/how-do-i-connect-copilot-to-a-genai-connection/299 "2026-02-04T17:35:19Z")

</div>

A: Once enabled, you must create a Copilot instance and link it to a Generative AI connection: Go to Administration \> Copilot and click Add Copilot. Select a Gen AI Connection (e.g., OpenAI or Azure OpenAI) from th…

---

## [How do I enable Copilot for my Sparkflows instance?](https://community.sparkflows.ai/t/how-do-i-enable-copilot-for-my-sparkflows-instance/298)

<div class="topic-metadata">

**Author:** [@Aayush](https://community.sparkflows.ai/u/Aayush)\
**Replies:** 0\
**Last updated:** [February 4, 2026, 5:31pm UTC](https://community.sparkflows.ai/t/how-do-i-enable-copilot-for-my-sparkflows-instance/298 "2026-02-04T17:31:13Z")

</div>

An administrator must first enable the feature in the platform settings: Navigate to Administration \> Configurations. Search for the setting module.enableCopilot and set its value to true. Click Save Configurati…

---

## [What are the main features of Copilot?](https://community.sparkflows.ai/t/what-are-the-main-features-of-copilot/297)

<div class="topic-metadata">

**Author:** [@Aayush](https://community.sparkflows.ai/u/Aayush)\
**Replies:** 0\
**Last updated:** [February 4, 2026, 5:29pm UTC](https://community.sparkflows.ai/t/what-are-the-main-features-of-copilot/297 "2026-02-04T17:29:50Z")

</div>

Sparkflows Copilot features include: Workflow/Pipeline Generation: Build entire flows or individual stages from a single prompt. Node-Level Assistance: Generate SQL queries or math expressions inside a specific nod…

---

## [What is Sparkflows Copilot and how can it help my data engineering tasks?](https://community.sparkflows.ai/t/what-is-sparkflows-copilot-and-how-can-it-help-my-data-engineering-tasks/296)

<div class="topic-metadata">

**Author:** [@Aayush](https://community.sparkflows.ai/u/Aayush)\
**Replies:** 0\
**Last updated:** [February 4, 2026, 5:28pm UTC](https://community.sparkflows.ai/t/what-is-sparkflows-copilot-and-how-can-it-help-my-data-engineering-tasks/296 "2026-02-04T17:28:19Z")

</div>

Sparkflows Copilot is a generative AI assistant integrated directly into the platform. It allows you to create and modify workflows, pipelines, and entire projects using natural language prompts instead of manually dragg…

---

## [How to Save Multiple Sheets into the Same Excel File](https://community.sparkflows.ai/t/how-to-save-multiple-sheets-into-the-same-excel-file/295)

<div class="topic-metadata">

**Author:** [@Aayush](https://community.sparkflows.ai/u/Aayush)\
**Replies:** 0\
**Last updated:** [February 4, 2026, 5:25pm UTC](https://community.sparkflows.ai/t/how-to-save-multiple-sheets-into-the-same-excel-file/295 "2026-02-04T17:25:01Z")

</div>

How do I configure the Advanced Save Excel Node to output multiple datasets into different sheets/tabs of a single Excel file? Solution To save multiple sheets to the same file without overwriting previous data, you mus…

---

## [Why does Tokenize to Rows keep null rows?](https://community.sparkflows.ai/t/why-does-tokenize-to-rows-keep-null-rows/294)

<div class="topic-metadata">

**Author:** [@Sanskar](https://community.sparkflows.ai/u/Sanskar)\
**Replies:** 0\
**Last updated:** [January 31, 2026, 6:24am UTC](https://community.sparkflows.ai/t/why-does-tokenize-to-rows-keep-null-rows/294 "2026-01-31T06:24:30Z")

</div>

Because row explosion should never drop data implicitly. If a row has: No matches Or a null input The node emits a row with null instead of removing it. Why? Dropping rows during tokenization can: Break ro…

---

## [Why does Lookup sometimes append nulls even when a match exists?](https://community.sparkflows.ai/t/why-does-lookup-sometimes-append-nulls-even-when-a-match-exists/292)

<div class="topic-metadata">

**Author:** [@Sanskar](https://community.sparkflows.ai/u/Sanskar)\
**Replies:** 0\
**Last updated:** [January 31, 2026, 6:21am UTC](https://community.sparkflows.ai/t/why-does-lookup-sometimes-append-nulls-even-when-a-match-exists/292 "2026-01-31T06:21:24Z")

</div>

Because first match wins. Append mode stops scanning lookup rows as soon as it finds a match. If the match occurs after earlier non-matches, that’s fine. But if: Whole-word is enabled Or match location is restri…

---

## [Why doesn’t Select Records preserve the original Excel row order?](https://community.sparkflows.ai/t/why-doesn-t-select-records-preserve-the-original-excel-row-order/291)

<div class="topic-metadata">

**Author:** [@Sanskar](https://community.sparkflows.ai/u/Sanskar)\
**Replies:** 0\
**Last updated:** [January 31, 2026, 6:20am UTC](https://community.sparkflows.ai/t/why-doesn-t-select-records-preserve-the-original-excel-row-order/291 "2026-01-31T06:20:52Z")

</div>

It actually does — but not the way you might expect. Spark DataFrames have no guaranteed row order. To make row selection deterministic, the node assigns an internal row number using: monotonically\_increasing\_id() …

---

## [When “Top N” and “After Record” are both set, which one wins?](https://community.sparkflows.ai/t/when-top-n-and-after-record-are-both-set-which-one-wins/290)

<div class="topic-metadata">

**Author:** [@Sanskar](https://community.sparkflows.ai/u/Sanskar)\
**Replies:** 0\
**Last updated:** [January 31, 2026, 6:19am UTC](https://community.sparkflows.ai/t/when-top-n-and-after-record-are-both-set-which-one-wins/290 "2026-01-31T06:19:51Z")

</div>

Short answer: neither “wins” — they’re merged. Behind the scenes, Select Records doesn’t treat these options as mutually exclusive. It normalizes everything into row ranges and applies them together. So: Top N = 10 …

---

## [Why does Lookup (Replace) sometimes not replace all matches?](https://community.sparkflows.ai/t/why-does-lookup-replace-sometimes-not-replace-all-matches/289)

<div class="topic-metadata">

**Author:** [@Sanskar](https://community.sparkflows.ai/u/Sanskar)\
**Replies:** 0\
**Last updated:** [January 31, 2026, 6:18am UTC](https://community.sparkflows.ai/t/why-does-lookup-replace-sometimes-not-replace-all-matches/289 "2026-01-31T06:18:41Z")

</div>

What users observe Only the first matching value is replaced Remaining matches stay unchanged Why this happens By default: replaceMultipleItems = false The node stops after the first successful match T…

---

## [Why does Select Records behave differently in preview vs execution?](https://community.sparkflows.ai/t/why-does-select-records-behave-differently-in-preview-vs-execution/288)

<div class="topic-metadata">

**Author:** [@Sanskar](https://community.sparkflows.ai/u/Sanskar)\
**Replies:** 0\
**Last updated:** [January 31, 2026, 6:17am UTC](https://community.sparkflows.ai/t/why-does-select-records-behave-differently-in-preview-vs-execution/288 "2026-01-31T06:17:52Z")

</div>

What users observe Fewer rows appear when running locally (preview mode) Full dataset appears during full workflow execution Why this happens Select Records applies a local safety limit: In local synchronous…

---

## [When should I use each algorithm available in sparkflows, how do they work, and what are the most important settings to tune?](https://community.sparkflows.ai/t/when-should-i-use-each-algorithm-available-in-sparkflows-how-do-they-work-and-what-are-the-most-important-settings-to-tune/180)

<div class="topic-metadata">

**Author:** [@shreyash\_prashu](https://community.sparkflows.ai/u/shreyash_prashu)\
**Replies:** 0\
**Last updated:** [December 20, 2025, 7:26pm UTC](https://community.sparkflows.ai/t/when-should-i-use-each-algorithm-available-in-sparkflows-how-do-they-work-and-what-are-the-most-important-settings-to-tune/180 "2025-12-20T19:26:54Z")

</div>

The platform offers a range of supervised and unsupervised machine learning algorithms, each optimized for different data types, business goals, and modeling constraints. The comparison below helps you understand when to…

---

## [How can I interpret individual predictions and get feature importance (Shapley values)?](https://community.sparkflows.ai/t/how-can-i-interpret-individual-predictions-and-get-feature-importance-shapley-values/179)

<div class="topic-metadata">

**Author:** [@shreyash\_prashu](https://community.sparkflows.ai/u/shreyash_prashu)\
**Replies:** 0\
**Last updated:** [December 19, 2025, 5:17pm UTC](https://community.sparkflows.ai/t/how-can-i-interpret-individual-predictions-and-get-feature-importance-shapley-values/179 "2025-12-19T17:17:08Z")

</div>

Sometimes it isn’t enough to know which features are important overall; you must be able to justify individual predictions. The H2O Scoring Node in Sparkflows enables this by calculating Shapley (SHAP) values at the row …

---

## [What are the most critical parameters to tweak for each specific H2O model type?](https://community.sparkflows.ai/t/what-are-the-most-critical-parameters-to-tweak-for-each-specific-h2o-model-type/178)

<div class="topic-metadata">

**Author:** [@shreyash\_prashu](https://community.sparkflows.ai/u/shreyash_prashu)\
**Replies:** 0\
**Last updated:** [December 19, 2025, 5:05pm UTC](https://community.sparkflows.ai/t/what-are-the-most-critical-parameters-to-tweak-for-each-specific-h2o-model-type/178 "2025-12-19T17:05:58Z")

</div>

To get strong performance from H2O models, it’s important to tune the parameters that matter most for each algorithm. Below is a practical guide to the key “levers” for each major H2O model type, based on common Sparkflo…

---

## [How can I create custom nodes in Java/Scala?](https://community.sparkflows.ai/t/how-can-i-create-custom-nodes-in-java-scala/116)

<div class="topic-metadata">

**Author:** [@Ragita](https://community.sparkflows.ai/u/Ragita)\
**Replies:** 0\
**Last updated:** [December 15, 2025, 1:14pm UTC](https://community.sparkflows.ai/t/how-can-i-create-custom-nodes-in-java-scala/116 "2025-12-15T13:14:58Z")

</div>

Sparkflows follows an open and extensible architecture, allowing developers to add new custom nodes/processors that can be exposed in Fire UI and embedded into workflows. The details for creating custom nodes are availa…

---

## [Where do I find the logs of the Sparkflows Server?](https://community.sparkflows.ai/t/where-do-i-find-the-logs-of-the-sparkflows-server/104)

<div class="topic-metadata">

**Author:** [@Tarika](https://community.sparkflows.ai/u/Tarika)\
**Replies:** 0\
**Last updated:** [December 15, 2025, 10:55am UTC](https://community.sparkflows.ai/t/where-do-i-find-the-logs-of-the-sparkflows-server/104 "2025-12-15T10:55:02Z")

</div>

The logs of the Sparkflows web server is in the file fireserver.log. Logs of the Sparkflows processes are in the file fire.log under the directory where Sparkflows has been installed.

---

## [What kinds of Data Sources are supported?](https://community.sparkflows.ai/t/what-kinds-of-data-sources-are-supported/103)

<div class="topic-metadata">

**Author:** [@Tarika](https://community.sparkflows.ai/u/Tarika)\
**Replies:** 0\
**Last updated:** [December 15, 2025, 10:22am UTC](https://community.sparkflows.ai/t/what-kinds-of-data-sources-are-supported/103 "2025-12-15T10:22:41Z")

</div>

Sparkflows supports reading data from a variety of Data Sources. These include: MySQL, JDBC, MongoDB, HBase, Cassandra, S3, HDFS, ADLS, DBFS, Oracle, Oracle NetSuite, TeraData, PostgreSQL, RedShift, SAP HANA, Salesforce…

---

## [When running on a Apache Spark cluster how does Sparkflows submit the spark jobs?](https://community.sparkflows.ai/t/when-running-on-a-apache-spark-cluster-how-does-sparkflows-submit-the-spark-jobs/88)

<div class="topic-metadata">

**Author:** [@Tarika](https://community.sparkflows.ai/u/Tarika)\
**Replies:** 1\
**Last updated:** [December 15, 2025, 4:41am UTC](https://community.sparkflows.ai/t/when-running-on-a-apache-spark-cluster-how-does-sparkflows-submit-the-spark-jobs/88 "2025-12-15T04:41:55Z")

</div>

Sparkflows uses spark-submit to submit the Apache Spark jobs to the cluster. Hence it is important that spark-submit works from the machine on which Sparkflows is installed.

---

## [How can I test the API of Sparkflows?](https://community.sparkflows.ai/t/how-can-i-test-the-api-of-sparkflows/90)

<div class="topic-metadata">

**Author:** [@Tarika](https://community.sparkflows.ai/u/Tarika)\
**Replies:** 0\
**Last updated:** [December 12, 2025, 2:20pm UTC](https://community.sparkflows.ai/t/how-can-i-test-the-api-of-sparkflows/90 "2025-12-12T14:20:53Z")

</div>

You can run below curl command: curl --location --request POST ‘http://localhost:8090/messageFromSparkJob’ --header ‘Content-Type: application/json’ --data-raw ‘{“jobId”: “256”, “message”: “this is test message”}’ Ex…

---

## [How does Apache Spark’s distributed execution affect the number of output files, and what method ensures saving the result in only one file?](https://community.sparkflows.ai/t/how-does-apache-spark-s-distributed-execution-affect-the-number-of-output-files-and-what-method-ensures-saving-the-result-in-only-one-file/89)

<div class="topic-metadata">

**Author:** [@Tarika](https://community.sparkflows.ai/u/Tarika)\
**Replies:** 0\
**Last updated:** [December 12, 2025, 2:13pm UTC](https://community.sparkflows.ai/t/how-does-apache-spark-s-distributed-execution-affect-the-number-of-output-files-and-what-method-ensures-saving-the-result-in-only-one-file/89 "2025-12-12T14:13:50Z")

</div>

Apache Spark runs distributed. As a result the data is partitioned across multiple process/machines. When any of the Save nodes is used to write the output to files, the number of files created is dependent on the numbe…

[Next page](https://community.sparkflows.ai/tag/faq/3.md?match_all_tags=true&page=1&tags%5B%5D=faq)
