Apache Spark SQL Parser for Java

General SQL Parser (GSP) parses Apache Spark SQL with a dedicated, hand-tuned grammar — full AST access, offline syntax validation, formatting, rewriting, and column-level data lineage, including complete procedural SQL support. One commercial SDK — Apache Spark SQL is available in the Java edition.

99%
of 70 documented Apache Spark SQL constructs parse fully — measured on GSP 4.1.11 (2026-08-10)
330
Apache Spark SQL keywords recognized by the grammar
327
keywords the parser also accepts as identifiers
71
Apache Spark SQL test SQL files in the regression corpus
Weekly
release cadence — dialect fixes ship in days

Apache Spark SQL that GSP parses

CACHE TABLE sales_cache AS
SELECT region, SUM(amt) AS total
FROM sales WHERE year = 2026
GROUP BY region;

Parse it in Java

import gudusoft.gsqlparser.*;

TGSqlParser parser = new TGSqlParser(EDbVendor.dbvsparksql);
parser.sqltext = sql; // the Apache Spark SQL above
if (parser.parse() == 0) {
    System.out.println(parser.sqlstatements.size() + " statement(s) parsed");
} else {
    System.out.println(parser.getErrormessage());
}

Java edition only. The .NET library has no Apache Spark SQL grammar, so this dialect cannot be parsed from C# or VB.NET. See the .NET SQL Parser SDK for the dialects it does cover.

What GSP handles in Apache Spark SQL

Beyond parsing, the same AST powers table/column extraction, validation, formatting, and column-level lineage for Apache Spark SQL.

Recent Apache Spark SQL parser updates

Full history in the release notes.

Apache Spark SQL resources

Common questions

Does GSP parse Apache Spark SQL stored procedures and procedural SQL?

Yes. GSP parses SQL scripting control flow: FOR loops and LOOP/END LOOP statements into a full AST — procedure bodies become real statement trees you can traverse, not opaque text blocks. This is what makes lineage and impact analysis work inside procedural code.

Do I need a live Apache Spark SQL database connection to validate SQL?

No. GSP validates Apache Spark SQL syntax completely offline. You get error line and column positions, the offending token, and a hint — with no server, driver, or credentials involved.

How complete is the Apache Spark SQL grammar?

GSP uses a dedicated hand-tuned grammar for Apache Spark SQL — not a generic SQL grammar with flags. The parser recognizes 330 Apache Spark SQL keywords and knows which 327 of them can also be used as identifiers, which is exactly the kind of edge case that breaks generic parsers.

Can I use the Apache Spark SQL parser from C# / .NET?

Not currently. Apache Spark SQL is a Java-edition dialect — the .NET library ships 15 dedicated grammars and Apache Spark SQL is not one of them, so there is no C# code path for it. Everything else on this page describes the Java SDK. If you need Apache Spark SQL in .NET, email info@sqlparser.com; customer demand is how dialects get prioritized for the .NET port.

Can GSP extract column-level lineage from Apache Spark SQL?

Yes. The built-in DataFlowAnalyzer produces column-level lineage, impact analysis, and call graphs from Apache Spark SQL scripts — it is the engine behind Gudu SQLFlow and the lineage integrations for DataHub and OpenMetadata.