Methodology - Processed Data

  • blank blank
Description:

Feature-selected regression and clustering tables, the accessibility_score target variable, and raw Weka output from the Project Methodology assignment.

Lee, Leone • Aug 10, 2026 08:41:47
Outputs of the Project Methodology assignment, building on dallas_last_mile_dataset.csv from the Data Preparation stage (Processed Data dataset). dallas_methodology_dataset.csv adds the new accessibility_score column (equity-weighted composite: 0.5·z(resource_density) − 0.25·z(poverty_rate) − 0.25·z(zero_vehicle_hh_pct)) and drops the redundant *_count columns. regression_table*.csv/.arff and clustering_table*.csv/.arff are feature-selected working tables (_selected versions have the final chosen features only) used for the regression and clustering models respectively. The *_result.txt files are raw Weka output (feature selection, clustering, classifier evaluation) documenting exact settings and results. cluster_assignments.csv gives each station's K-Means and EM cluster label alongside its accessibility_score. pca_loadings.csv gives PCA component loadings from the unsupervised feature selection step. Software used: Weka 3.8.6 (command-line and GUI) and Python 3 (scikit-learn, pandas). Same 46-station Dallas DART walkshed study area as Data Preparation.

Metadata

Name Value Last Modified
Time Periods
  • Same underlying data as Data Preparation (see above); analysis performed August 2026
Added by Lee, Leone on Aug 10, 2026
  • Time Periods: Same underlying data as Data Preparation (see above); analysis performed August 2026

No extraction events recorded.

Statistics

Views: 162
Last viewed: Sep 26, 2026 22:24:34
Downloads: 0
Last downloaded: Never
Last Modified: Aug 10, 2026 08:51:45

Spaces containing the Dataset

7 datasets |

Collections containing the Dataset

3 datasets |

Tags