| FazBrowse GitHub Viewer | Trending | | Home |
| Tools: [Download Repo ZIP] [Original HTTPS Page] |
| Name | Name | Last commit date | ||
|---|---|---|---|---|
AWARE Narrator is a comprehensive Python toolkit that processes sensor data from mobile devices, performs DBSCAN clustering on location data, and generates detailed narrative descriptions of user mobility and activity patterns. The toolkit integrates multiple sensors including location, applications, keyboard input, screen usage, calls, messages, and more. It includes Google Maps API integration for reverse geocoding and uses a configuration file (config.yaml) to customize parameters.
If you plan to use Google Maps API for reverse geocoding (requires USE_GOOGLE_MAP: true and a valid API key file referenced by GOOGLE_MAP_KEY in config.yaml), you may need to manually update the geocoding module due to recent changes in the Google Maps Services Python library:
Alternative: A copy of the updated geocoding.py file has been included in this project for convenience. You can copy it directly to replace the installed package in your current environment:
# Make sure you're in the correct environment first
mamba activate my_env # or conda activate my_env
# Find your googlemaps package location in the current environment
python -c "import googlemaps; print(googlemaps.__file__)"
# Copy the included geocoding.py to replace the installed version
cp geocoding.py $(python -c "import googlemaps; import os; print(os.path.dirname(googlemaps.__file__))")/geocoding.pyThis manual update ensures compatibility with the address descriptor feature in geocoding.py
Note: This setup is only required if you plan to use Google Maps API for reverse geocoding. If you set USE_GOOGLE_MAP: false in your config.yaml, the toolkit will work without this update.
This project supports Mamba and Conda for managing and installing dependencies. We recommend using Mamba for faster package resolution and installation.
Ensure you have Mamba installed.
Activate your Mamba environment (or create one if needed):
mamba create -n my_env python=3.13
mamba activate my_envCreate (or update) your environment from the file:
# Create a new environment with the name defined in the file:
mamba env create --file environment.yml
# Or create under a custom name:
mamba env create -n <my_env> --file environment.yml
# To update an existing environment to match environment.yml:
mamba env update -n <my_env> --file environment.yml --pruneActivate your Conda environment (or create one if needed):
conda create -n my_env python=3.13
conda activate my_envCreate (or update) your environment from the file:
# Create a new environment with the name defined in the file:
conda env create --file environment.yml
# Or create under a custom name:
conda env create -n <my_env> --file environment.yml
# To update an existing environment to match environment.yml:
conda env update -n <mv_env> --file environment.yml --pruneNote: environment.yml was generated with:
conda env export --from-history > environment.yml
The project includes several Python scripts:
The scripts require a YAML configuration file with the following structure:
# Configuration for Aware Narrator
# Two modes: manual or auto
# Manual mode: Uses P_IDs, START_TIME, and END_TIME, output_file
# Auto mode: Uses survey_time_file, time_ranges, output_dir
MODE: "manual" # Options: "manual" or "auto"
# START of manual mode configuration (used when MODE: "manual")
P_IDs:
- SS001
START_TIME: "2025-05-25 06:00:00"
END_TIME: "2025-06-04 23:59:00"
output_file: "description/{P_ID}_{START_TIME}_{END_TIME}.txt"
# END of manual mode configuration
# START of auto mode configuration (used when MODE: "auto")
# CSV File containing survey-specific answered timestamp, one pid might have multiple rows
# Main columns: pid,survey_id,survey_time_unix
survey_time_file: "resources/survey_time.csv"
# Direction for time range processing:
# - "backward": survey_time is the END point, process data BEFORE survey (default, for weekly surveys)
# e.g., 7d range = [survey_time - 7d, survey_time]
# - "forward": survey_time is the START point, process data AFTER that time (for cumulative from baseline)
# e.g., 7d range = [survey_time, survey_time + 7d]
direction: "backward" # Options: "backward" or "forward"
# Alignment for the reference timestamp:
# - true: Align to midnight (00:00:00) of the survey timestamp's day
# - false: Use the exact survey timestamp as-is (default)
align_to_midnight: false
# Skip records from the survey date itself:
# - true: Exclude all data from the calendar day of the survey
# - false: Include survey day data (default)
skip_survey_day: false
# Time ranges to save descriptions
# Option 1: Specify start and end to auto-generate all ranges in between (same unit required)
# Order follows start→end (ascending if start < end, descending if start > end)
time_range_start: 7d
time_range_end: 1d
#
# Option 2: Explicit list (used if time_range_start/time_range_end are not set)
# time_ranges:
# - 7d
# - 3d
# - 1d
output_dir: "description/{P_ID}"
# END of auto mode configuration
# Replace by your own mapping csv.
# Must contain device_id and pid columns.
# Multiple device ids need to be splitted by ";"
# For example:
# Header: pid,device_id
# Row 1: 1234,aaaa-bbbb-cccc;1111-082a-4a73-8ee3
# Row 2: SS11,fa1da-3adrv-123a
pid_to_deviceid_map: "resources/pid_deviceid_mapping.csv" # generated by running map_pid_deviceid.py. See README for instructions.
timezone: "Australia/Melbourne" # replace by actual timezone
input_directory: "participant_data/{P_ID}"
package_to_app_map: "resources/app_package_pairs.jsonl" # required; generated by the screentext preprocess pipeline
session_data_file: "step1_data/{P_ID}/sessions.jsonl" # required for applicaion and keyboard analysis; generated by running extract_sessions.py with screen.jsonl for the corresponding P_ID
cleaned_screentext_file: "step1_data/{P_ID}/clean_input.jsonl" # required for screen text description; generated by using screen text preprocessing pipeline
reverse_geocoding_output_dir: "locations_query_results/{P_ID}"
daily_output_dir: "daily_description/{P_ID}"
num_workers: 1 # Number of parallel workers for processing participants (1 = sequential, >1 = parallel via ProcessPoolExecutor)
sensor_integration_time_window: 60 # minutes
gate_time_window: 5 # minutes; required for wifi and bluetooth scan data integration.
sensors:
- "applications_foreground"
- "applications_notifications"
- "battery"
- "bluetooth"
- "calls"
- "installations"
- "keyboard"
- "messages"
- "screen"
- "screentext"
- "wifi"
- "sensor_wifi"
- "locations"
DISCARD_SYSTEM_UI: true # Applied to 'applications_foreground', 'applications_notifications', 'installations' based on system_ui_apps
USE_GOOGLE_MAP: false # Set to true to enable Google Maps reverse geocoding
GOOGLE_MAP_KEY: "GOOGLE_MAP_API_KEY.txt" # Path to a text file containing your Google API key (only read when USE_GOOGLE_MAP is true)
eps: 0.000047 # DBSCAN clustering parameter: radians (0.000047 radians × 6371000 m ≈ 300m)
min_samples: 3 # DBSCAN clustering parameter: mininum number of points to form a cluster
location_minimum_data_points: 3 # Minimum number of location data points to display a place in location description
location_minimum_stay_minutes: 3 # Minimum stay duration in minutes to display a place in location description
night_time_start: 22 # Start of nighttime in 24-hour format, used for determining home location
night_time_end: 6 # End of nighttime in 24-hour format, used for determining home location
merge_distance_threshold: 300 # Distance threshold in meters to merge home candidates and clusters with no night points
# Using package instead of app name in case of locales (different languages for the same app name)
blacklist_apps:
- com.aware.phone # AWARE-Light
system_ui_apps:
- com.android.systemui # System UI
- com.sec.android.app.launcher # One UI Home / One UI 首頁 / One UI 主屏幕 / TouchWiz home / Samsung Experience Home / Écran d'accueil One UI
- com.samsung.android.app.cocktailbarservice # Edge panels
- com.huawei.android.launcher # 华为桌面 / Huawei Home / Beranda Huawei
- com.miui.home # System launcher / Peluncur sistem / 系统桌面
- com.oppo.launcher # System Launcher / システムランチャー
- com.google.android.apps.nexuslauncher # Pixel Launcher / Peluncur Pixel / Lanceur d'applications Pixel
- com.motorola.launcher3 # Moto App Launcher
- net.oneplus.launcher # OnePlus Launcher
- jp.co.sharp.android.launcher3 # AQUOS Home
- com.android.launcher3 # Quickstep
- com.android.launcher # System Launcher
- com.vivo.hiboard # Jovi Home
- com.mi.android.globallauncher # POCO Launcher
- com.bbk.launcher2 # System launcher / 系统桌面
- com.sec.android.app.desktoplauncher # Samsung DeX home
- com.sec.android.emergencylauncher # Launcher
- com.hihonor.android.launcher # 荣耀桌面 / HONOR Home
- com.sonymobile.launcher # Xperia主屏幕 / Xperia Home
- com.google.android.inputmethod.latin # Gboard
Generate the participant ID to device ID mapping file required by the main analysis script:
python map_pid_deviceid.pyThis script processes a CSV file containing participant information and creates a mapping file used by aware_narrator.py.
Input Requirements:
Default behavior:
The output mapping file is referenced in config.yaml as pid_to_deviceid_map and is required for the main analysis script to function properly.
Split JSONL files from the exported data into participant-specific directories:
# Split all JSONL files for all participants
python split_by_participant.py
# Split specific JSONL files only
python split_by_participant.py --jsonl-files locations applications_foreground
# Process only specific participants
python split_by_participant.py --pids SS001 SS002 SS003
# Run threshold analysis mode to analyze data distribution
python split_by_participant.py --threshold-analysis
# Custom input/output directories
python split_by_participant.py --input-dir exported_jsonl --output-dir participant_dataThis script processes JSONL files from the exported_jsonl directory and splits them into the participant_data/{P_ID}/ structure required by the main analysis.
Features:
Process screentext data through the complete preprocessing pipeline:
# Process a single participant
python run_screentext_preprocess_pipeline.py --participant SS001
# Process all participants (Step 1 sequential, Steps 2-5 parallel)
python run_screentext_preprocess_pipeline.py --all
# Process specific participants only
python run_screentext_preprocess_pipeline.py --all --include SS001 SS002 SS003
# Process all participants except specific ones
python run_screentext_preprocess_pipeline.py --all --exclude SS001 SS002
# Custom timezone and worker threads
python run_screentext_preprocess_pipeline.py --all --timezone "Australia/Melbourne" --workers 8Pipeline Steps:
Note: The screentext preprocessing pipeline includes session extraction functionality, so you do NOT need to run extract_sessions.py separately if you're using the screentext pipeline. The pipeline generates the required sessions.jsonl file as part of Step 5.
If you're not using the screentext preprocessing pipeline, you can extract session boundaries separately:
# Process a single participant
python extract_sessions.py --participant SS001
# Process all participants in the input directory
python extract_sessions.py --all
# Custom session threshold (default: 45000ms = 45 seconds)
python extract_sessions.py --participant SS001 --threshold 45000
# Custom input/output directories
python extract_sessions.py --participant SS001 --input-dir custom_data --output-dir custom_outputImportant: Only run this script if you're NOT using the screentext preprocessing pipeline, as the pipeline already includes session extraction.
Convert JSON files to JSONL format:
# Convert all JSON files in a folder
python json2jsonl.py /path/to/json/folder
# Specify output folder
python json2jsonl.py /path/to/json/folder -o /path/to/output/folder
# Search recursively in subdirectories
python json2jsonl.py /path/to/json/folder -r -o /path/to/output/folderpython aware_narrator.pyThis processes all sensor data according to the configuration and generates comprehensive narratives.
Split output narratives by sensor type:
python split_description.pyThis creates separate files for each sensor type in the description_split/{PID}/ folder.
The project expects the following directory structure:
exported_jsonl/ # Raw exported JSONL files (for split_by_participant.py)
├── applications_foreground.jsonl
├── applications_notifications.jsonl
├── battery.jsonl
├── bluetooth.jsonl
├── calls.jsonl
├── installations.jsonl
├── keyboard.jsonl
├── locations.jsonl
├── messages.jsonl
├── screen.jsonl
├── screentext.jsonl # Required for screentext analysis
├── wifi.jsonl
└── sensor_wifi.jsonl
participant_data/ # After running split_by_participant.py
├── {P_ID}/
│ ├── applications_foreground.jsonl
│ ├── applications_notifications.jsonl
│ ├── battery.jsonl
│ ├── bluetooth.jsonl
│ ├── calls.jsonl
│ ├── installations.jsonl
│ ├── keyboard.jsonl
│ ├── locations.jsonl
│ ├── messages.jsonl
│ ├── screen.jsonl
│ ├── screentext.jsonl # Required for screentext analysis
│ ├── wifi.jsonl
│ └── sensor_wifi.jsonl
step1_data/ # After running screentext pipeline or extract_sessions.py
├── {P_ID}/
│ ├── sessions.jsonl # Generated by screentext pipeline or extract_sessions.py
│ └── clean_input.jsonl # Generated by screentext pipeline
resources/
├── pid_deviceid_mapping.csv # Generated by map_pid_deviceid.py
└── app_package_pairs.jsonl # Generated by screentext pipeline
The toolkit generates several types of output:
The toolkit analyzes the following sensor types:
This project is for research purposes. Contact the developers for usage permissions.
For questions, reach out to the maintainers of the AWARE Narrator project.
| Back | FazBrowse Home | New Git URL |