This project does NS1 Log analysis using statistics and ML.
- Linux shell + some RAM
- Java 17 for Spark (Java 21+ won't work)
- Spark 3.4+ and its required Java version
- Python 3.12+ (required by
full_analysis.sh) - Pip packages from
requirements.txt
Use these when running locally on macOS/Linux:
export SPARK_LOCAL_IP=127.0.0.1
export JAVA_HOME="$(/usr/libexec/java_home -v 17)"Why these are needed:
SPARK_LOCAL_IPForces Spark to bind to localhost. This avoids hostname/network interface resolution problems (common on laptops and VPN setups) that can cause Spark startup failures or flaky local runs.JAVA_HOMEPins Spark to Java 17 explicitly. Spark in this project is expected to run with Java 17; pointing to newer JVMs can fail at startup or runtime.PYTHON_LOGGING_LEVEL(optional) Controls Python logging verbosity in the analysis scripts viaea_common.py. Example:
export PYTHON_LOGGING_LEVEL=DEBUGAccepted values follow standard Python logging levels, for example DEBUG, INFO, WARNING, ERROR, CRITICAL.
For full-run, execute in a *nix terminal:
bash ./full_analysis.shLogs under data/all_logs/ may be arranged in subdirectories, such as by
server and year. preprocess.sh preserves those relative paths through
data/pre-processed1/ and data/pre-processed2/, and log_parser.py
processes them recursively. This allows separate logs with the same filename
to be imported without renaming them:
data/all_logs/server-a/2024/L0101000.log
data/all_logs/server-b/2025/L0101000.log
The parser's incremental state tracks the relative path, so these are treated
as distinct inputs. Parsed log_files.filename remains the original basename;
its ID and SHA-256 hash distinguish records in the output data.