Apache Spark SQL Parser for Java
General SQL Parser (GSP) parses Apache Spark SQL with a dedicated, hand-tuned grammar — full AST access, offline syntax validation, formatting, rewriting, and column-level data lineage, including complete procedural SQL support. One commercial SDK — Apache Spark SQL is available in the Java edition.
Apache Spark SQL that GSP parses
CACHE TABLE sales_cache AS
SELECT region, SUM(amt) AS total
FROM sales WHERE year = 2026
GROUP BY region; Parse it in Java
import gudusoft.gsqlparser.*;
TGSqlParser parser = new TGSqlParser(EDbVendor.dbvsparksql);
parser.sqltext = sql; // the Apache Spark SQL above
if (parser.parse() == 0) {
System.out.println(parser.sqlstatements.size() + " statement(s) parsed");
} else {
System.out.println(parser.getErrormessage());
} Java edition only. The .NET library has no Apache Spark SQL grammar, so this dialect cannot be parsed from C# or VB.NET. See the .NET SQL Parser SDK for the dialects it does cover.
What GSP handles in Apache Spark SQL
- CACHE TABLE
- REFRESH TABLE
- ADD JAR/FILE resources
- explode(from_json(...)) lineage
- FOR and LOOP statements
- Procedural SQL: SQL scripting control flow: FOR loops and LOOP/END LOOP statements
Beyond parsing, the same AST powers table/column extraction, validation, formatting, and column-level lineage for Apache Spark SQL.
Recent Apache Spark SQL parser updates
- 2026-06-02 [SparkSQL] Parse PRIMARY KEY / FOREIGN KEY in CREATE TABLE
- 2026-05-29 [Hive, SparkSQL] Structured-dataflow lineage for explode(from_json(...))
- 2026-03-03 [SparkSQL] Support FOR loop statement
Full history in the release notes.
Apache Spark SQL resources
- Apache Spark SQL syntax reference — dialect documentation
- Apache Spark SQL keyword compatibility — every keyword, and whether it can be an identifier
- Apache Spark SQL capability data (JSON) — measured construct-by-construct parse coverage, machine-readable
- Java quick start — Maven/Gradle setup and first parse
Common questions
Does GSP parse Apache Spark SQL stored procedures and procedural SQL?
Yes. GSP parses SQL scripting control flow: FOR loops and LOOP/END LOOP statements into a full AST — procedure bodies become real statement trees you can traverse, not opaque text blocks. This is what makes lineage and impact analysis work inside procedural code.
Do I need a live Apache Spark SQL database connection to validate SQL?
No. GSP validates Apache Spark SQL syntax completely offline. You get error line and column positions, the offending token, and a hint — with no server, driver, or credentials involved.
How complete is the Apache Spark SQL grammar?
GSP uses a dedicated hand-tuned grammar for Apache Spark SQL — not a generic SQL grammar with flags. The parser recognizes 330 Apache Spark SQL keywords and knows which 327 of them can also be used as identifiers, which is exactly the kind of edge case that breaks generic parsers.
Can I use the Apache Spark SQL parser from C# / .NET?
Not currently. Apache Spark SQL is a Java-edition dialect — the .NET library ships 15 dedicated grammars and Apache Spark SQL is not one of them, so there is no C# code path for it. Everything else on this page describes the Java SDK. If you need Apache Spark SQL in .NET, email info@sqlparser.com; customer demand is how dialects get prioritized for the .NET port.
Can GSP extract column-level lineage from Apache Spark SQL?
Yes. The built-in DataFlowAnalyzer produces column-level lineage, impact analysis, and call graphs from Apache Spark SQL scripts — it is the engine behind Gudu SQLFlow and the lineage integrations for DataHub and OpenMetadata.