FazBrowse GitHub Viewer | Trending |
URL:
| Home
Tools: [Download Repo ZIP]   [View Raw Code]   [Original HTTPS Page]

OneCite/docs/python_api.rst at main · HzaCode/OneCite · GitHub

Latest commit

 

History

History
361 lines (262 loc) · 9.36 KB

File metadata and controls

361 lines (262 loc) · 9.36 KB

Python API Reference

This guide covers using OneCite as a Python library in your own code.

Basic Usage

Simple Citation Processing

from onecite import process_references

# Process a simple reference
result = process_references(
    input_content="10.1038/nature14539",
    input_type="txt",
    template_name="journal_article_full",
    output_format="bibtex"
)

# Print results
for citation in result['results']:
    print(citation)

The Result Dictionary

The process_references function returns a dictionary containing:

  • results (List[str]): List of formatted citation strings, one per unique resolved work
  • report (dict): Processing report with keys:
    • total (int): Total number of entries processed
    • succeeded (int): Number of successfully processed entries
    • failed_entries (List[Dict]): Failed entries, each with the original input text (raw_text) and a reason code
    • warnings (List[Dict]): Non-blocking review warnings, e.g. text_metadata_mismatch when the input text appears to describe a different work than the resolved DOI
    • duplicates (List[Dict]): Entries whose DOI already resolved earlier in the batch
result = process_references(
    input_content="10.1038/nature14539",
    input_type="txt",
    template_name="journal_article_full",
    output_format="bibtex"
)

print(f"Total: {result['report']['total']}")
print(f"Succeeded: {result['report']['succeeded']}")
print(f"Failed: {len(result['report']['failed_entries'])}")

Processing Different Input Formats

Plain Text Input

from onecite import process_references

txt_content = """
10.1038/nature14539

arXiv:1706.03762

ISBN:9780262035613
"""

result = process_references(
    input_content=txt_content,
    input_type="txt",
    template_name="journal_article_full",
    output_format="bibtex"
)

# Access results
print('\n\n'.join(result['results']))

BibTeX Input

from onecite import process_references

bibtex_content = """
@article{LeCun2015,
    title = {Deep Learning},
    author = {LeCun, Yann and Bengio, Yoshua and Hinton, Geoffrey},
    journal = {Nature},
    year = {2015}
}
"""

result = process_references(
    input_content=bibtex_content,
    input_type="bib",
    template_name="journal_article_full",
    output_format="bibtex"
)

print('\n\n'.join(result['results']))

Output Formats

OneCite supports BibTeX and CSL-JSON output:

# BibTeX (default)
result = process_references(
    input_content="10.1038/nature14539",
    input_type="txt",
    template_name="journal_article_full",
    output_format="bibtex"
)

# CSL-JSON — each result string is one CSL item (JSON object),
# ready for pandoc, Quarto, citeproc, or reference-manager import
result = process_references(
    input_content="10.1038/nature14539",
    input_type="txt",
    template_name="journal_article_full",
    output_format="csl-json"
)

Passing any other value (such as "apa" or "mla") raises FormatError; styled rendering belongs to dedicated tools like pandoc or citeproc-py, which consume the CSL-JSON that OneCite emits.

Candidate Suggestions for Ambiguous References

For plain-text title searches, use the suggestion API instead of resolving directly to BibTeX:

from onecite import suggest_references

result = suggest_references(
    input_content="Deep learning Hinton",
    input_type="txt",
    limit=5,
)

for suggestion in result['suggestions']:
    for candidate in suggestion['candidates']:
        print(candidate['title'], candidate.get('doi', ''))

A note on interactive_callback

process_references accepts an interactive_callback keyword for backward compatibility, but it is never invoked: process is strictly non-interactive and fail-closed. Entries without a verifiable strong identifier stay unresolved and are reported. To review candidates for an ambiguous reference and then resolve one, take the DOI from the chosen suggestion and pass it back through process_references.

Advanced Data Structures

OneCite defines three TypedDict classes representing different stages of the processing pipeline:

RawEntry

A TypedDict representing an unprocessed reference entry (Stage 1):

from onecite import RawEntry
from typing import Dict, Any, Optional

# RawEntry is a TypedDict with these fields:
entry: RawEntry = {
    'id': 1,
    'raw_text': "10.1038/nature14539",
    'doi': "10.1038/nature14539",
    'url': None,
    'query_string': None,
    'original_entry': None
}

IdentifiedEntry

A TypedDict representing an entry after identification from data sources (Stage 2):

from onecite import IdentifiedEntry

# IdentifiedEntry includes fields like:
# id, raw_text, doi, arxiv_id, url, metadata, status

CompletedEntry

A TypedDict representing a completed or failed Stage-3 entry. A completed entry contains the metadata available to the formatter; it is not guaranteed to have every possible field:

from onecite import CompletedEntry

# CompletedEntry includes fields like:
# id, doi, status, bib_key, bib_data

These are TypedDict classes used internally by the pipeline. For typical usage, call process_references() directly.

Working with Templates

Load and inspect templates:

from onecite import TemplateLoader

loader = TemplateLoader()

# Load a specific template
template = loader.load_template("journal_article_full")
print(f"Template name: {template['name']}")
print(f"Entry type: {template['entry_type']}")
print(f"Fields: {[f['name'] for f in template['fields']]}")

# Use a custom templates directory
custom_loader = TemplateLoader(templates_dir="/path/to/templates")
custom_template = custom_loader.load_template("my_template")

Using the Pipeline Controller

For advanced use cases requiring more control over the processing pipeline:

from onecite import PipelineController

# Create controller (optionally enable Google Scholar)
controller = PipelineController(use_google_scholar=False)

# Process with full control
result = controller.process(
    input_content="10.1038/nature14539",
    input_type="txt",
    template_name="journal_article_full",
    output_format="bibtex"
)

print('\n\n'.join(result['results']))

For typical usage, process_references() is simpler and covers the same functionality.

Error Handling

Handling Exceptions

from onecite import process_references, ValidationError, ParseError

try:
    result = process_references(
        input_content="invalid_reference",
        input_type="txt",
        template_name="journal_article_full",
        output_format="bibtex"
    )
except ValidationError as e:
    print(f"Validation error: {e}")
except ParseError as e:
    print(f"Parse error: {e}")
except Exception as e:
    print(f"Processing error: {e}")

Processing Files

Reading from File

from onecite import process_references

# Read from file
with open("references.txt", "r", encoding="utf-8") as f:
    content = f.read()

result = process_references(
    input_content=content,
    input_type="txt",
    template_name="journal_article_full",
    output_format="bibtex"
)

# Write to file
output_content = '\n\n'.join(result['results'])
with open("output.bib", "w", encoding="utf-8") as f:
    f.write(output_content)

Complete Example

from onecite import process_references

# Read references
with open("my_references.txt", "r", encoding="utf-8") as f:
    references = f.read()

# Process with BibTeX format
result = process_references(
    input_content=references,
    input_type="txt",
    template_name="journal_article_full",
    output_format="bibtex"
)

# Check results
report = result['report']
print(f"Total entries: {report['total']}")
print(f"Successfully processed: {report['succeeded']}")
print(f"Failed: {len(report['failed_entries'])}")

if report['failed_entries']:
    print("\nFailed entries:")
    for failed in report['failed_entries']:
        print(f"  - Entry {failed['id']}: {failed.get('error', 'Unknown error')}")

# Save output
output_content = '\n\n'.join(result['results'])
with open("formatted_refs.txt", "w", encoding="utf-8") as f:
    f.write(output_content)

print("\nDone!")

API Reference

See :doc:`api/core` for the complete API documentation.

Next Steps


Back | FazBrowse Home | New Git URL