What a score cannot answer
PsyPredict shows why the target, preprocessing and pipeline state are part of the evaluation.
The PsyPredict notebooks compare Random Forest and logistic regression on a workplace mental-health survey. The target is the treatment response. Predicting that response is not equivalent to diagnosing someone or estimating their need for treatment.
A baseline needs comparable conditions
Logistic regression emits a convergence warning and has no scaling. Random Forest scores higher in the saved output, but that difference does not isolate the effect of algorithm choice. Data treatment is also being compared.
A saved artifact is not reproduction
Training and parameter-search outputs are versioned. The CSVs needed to run them are not. The SHAP step ends in a shape error; repository interfaces still address telco churn, a different domain.
The next useful question is not how to present a higher score, but how to produce a coherent data, training and evaluation workflow. The technical case study collects the sources and observed limitations.