API reference

ConfoState: membrane protein conformational state classification.

confostate.extract_features(pdb_path: str, pdb_id: str | None = None, family: str = 'LeuT', annotations_row: dict[str, Any] | None = None, reference_dir: str | None = None, include_rmsd: bool = True) → dict[str, float][source]

Extract all structural features from a PDB file.

Loads the structure once via MDAnalysis, then runs each feature module.

confostate.evaluate_model(model, X_test, y_test, labels: list[str] | None = None) → dict[source]

Compute standard classification metrics.

Returns a serializable dictionary.

confostate.get_baseline_models(random_state: int = 42) → dict[str, object][source]

Return baseline estimators keyed by model name.

confostate.get_registered_model(registry_path: str, family: str) → dict[str, Any] | None[source]

Return latest model entry for a family.

confostate.load_annotations(csv_path: str, family: str | None = None) → DataFrame[source]

Load structure annotations from a CSV file.

Parameters:
  • csv_path (str) – Path to the CSV file containing annotations.

  • family (str, optional) – Filter by protein family if provided.

Returns:

DataFrame with annotation columns from the CSV file.

Return type:

pd.DataFrame

Raises:
confostate.register_model(registry_path: str, family: str, model_name: str, artifact_path: str, metrics: dict[str, Any] | None = None, data_version: str | None = None, extra: dict[str, Any] | None = None) → dict[str, Any][source]

Append a model record to registry and return created record.

confostate.run_training(annotations_csv: str, features_csv: str, model_name: str, out_dir: str, family: str | None = None, test_size: float = 0.2, random_state: int = 42) → dict[str, Any][source]

Run a train/eval cycle for a selected baseline model.

confostate.write_evaluation_report(metrics: dict, output_path: str, title: str = 'ConfoState Evaluation Report') → None[source]

Write a compact markdown report from metrics.

CSV data loader for conformational state annotations.

confostate.data.loader.load_annotations(csv_path: str, family: str | None = None) → DataFrame[source]

Load structure annotations from a CSV file.

Parameters:
  • csv_path (str) – Path to the CSV file containing annotations.

  • family (str, optional) – Filter by protein family if provided.

Returns:

DataFrame with annotation columns from the CSV file.

Return type:

pd.DataFrame

Raises:
confostate.data.loader.load_from_input_dir(input_dir: str = './input') → DataFrame[source]

Scan input directory for PDB files and return metadata.

Parameters:

input_dir (str) – Path to directory containing .pdb files.

Returns:

DataFrame with pdb_id and file_path for each .pdb file found.

Return type:

pd.DataFrame