Explosion 💥 @explosion.ai · 16/09/2024🔥 New case study: How GitLab built scalable spaCy pipelines to process a year's worth of support tickets and create actionable insights to better support their community. explosion.ai/blog/gitlab-... 122
Explosion 💥 @explosion.ai · 21/06/2024📝 Out now: How S&P Global shipped NLP pipelines for real-time commodities trading insights in a high-security environment with LLMs in the loop. 10× speed-up of their data workflows and up to 99% accuracy at 6mb! explosion.ai/blog/sp-glob... 010
Explosion 💥 @explosion.ai · 22/04/2024Out now: Thinc v9.0! 🔮 This release is the foundation of the upcoming spaCy v4 release and adds support for more powerful learning rates. We have also merged thinc-apple-ops into Thinc, so Apple AMX is supported out-of-the-box. Details & release notes: github.com/explosion/th... 032
Explosion 💥 @explosion.ai · 11/04/20245️⃣ Error analysis To maximize ROI from your data engineering, evaluation metrics should be paired with quantitative error analysis. Our latest example error analysis recipe iterates through false positives/negatives and lets you record the reasons to inform your improvement plan. 110
Explosion 💥 @explosion.ai · 11/04/20243️⃣ Model training During the training process, we recommend running Prodigy's train-curve command, which is a great way to quickly see whether more data of similar quality as the current dataset would improve the model. 110
Explosion 💥 @explosion.ai · 11/04/20242️⃣ Review Quantitative measurements of disagreements should always be accompanied by a qualitative analysis. Prodigy's review recipe is an excellent tool for that. We use it in all our consulting projects to inform and illustrate data model discussions: explosion.ai/tailored-sol... 110
Explosion 💥 @explosion.ai · 11/04/2024Prodigy provides built-in inter-annotator agreement commands (for tokens and text-level annotations) that you can run directly on your annotated dataset. It also lets you configure custom overlap by specifying the expected number of annotations per example. 110
Explosion 💥 @explosion.ai · 11/04/20241️⃣ Dataset development Data development is an iterative process. It’s good practice to test your initial annotation scheme and guidelines during the pilot phase and measure the inter-annotator agreement. 110
Explosion 💥 @explosion.ai · 27/03/2024New Prodigy plugin: prodigy-evaluate! 📈 confusion matrix and per-label stats 🔎 explore examples your model struggles with most 🍬 entity-level insights for NER with MantisNLP's nervaluate library github.com/explosion/pr... 033
Explosion 💥 @explosion.ai · 19/02/2024The SSO plugin is compatible with Prodigy >=1.15.0 and is part of our expanded company license offering, which also includes priority community and email support. 120
Explosion 💥 @explosion.ai · 19/02/2024Some data development and annotation projects need top-notch security. 🔒 Introducing the Prodigy Single Sign-On (SSO) plugin It's the first in a series of premium Prodigy plugins for company licenses. 122
Explosion 💥 @explosion.ai · 05/02/2024🗺️ Custom mapping: Instead of using a large skills taxonomy as the NER label scheme, generic skill entities are mapped onto a taxonomy using semantic similarity. Great work from the Nesta team & thanks to ESCoE for funding! 020
Explosion 💥 @explosion.ai · 05/02/2024✂️ Multiskill splitting: Nesta uses spaCy's dependency parsing to split up multi-skill phrases, like "developing apps and visualizations" into individual skills "developing apps" and "developing visualizations". 120
Explosion 💥 @explosion.ai · 05/02/2024New case study: How the Nesta data science team extracts skills from millions of online job ads to better understand UK skill demand, using spaCy and Prodigy. A few project highlights in this thread 🧵✨ explosion.ai/blog/nesta-s... 132
Explosion 💥 @explosion.ai · 25/01/2024The new entity linking functionality lets you to specify a KB and candidate selector. The LLM will then pick the most likely candidate, given the context. spacy.io/api/large-la... 030
Explosion 💥 @explosion.ai · 25/01/2024Out now: spacy-llm v0.7.0! 🔗 Built-in entity linking support 💬 New task for translation from/to arbitrary languages ❓ Use the Doc as prompt for question answering 🧩 Arbitrarily long docs via sharding github.com/explosion/sp... 185
Explosion 💥 @explosion.ai · 30/11/2023💌 OUT NOW: The latest edition of our spaCy newsletter featuring our new Merch Store, spaCy 3.7 and spacy-llm 0.6 releases, links to our latest talks, Nesta's Skills Extractor library, and new Prodi.gy blog on 2023 updates! 🚀 Read & sign up: us12.campaign-archive.com?u=83b0498b1e... 050
Explosion 💥 @explosion.ai · 30/10/2023Dealing with a huge bucket of images that you want to annotate? The new image retrieval features in Prodigy-ANN might help! To help explain this new feature, @koaning.bsky.social made a small demo to highlight the new feature 👀 youtu.be/vhbyekSsG8o 110
Explosion 💥 @explosion.ai · 26/10/2023🛠️ Improved DX of working with custom CS/JSS by supporting loading from local dirs and remote URLs Incorporate frameworks like HTMX for a dynamic interface using our latest Custom Events. prodi.gy/docs/custom-.... 100
Explosion 💥 @explosion.ai · 26/10/2023💥 Prodigy 1.14.5 is out! 💥 We've focused on the front end 💅 prodi.gy/docs/changelog ☑️ A new toggle between token and character-based highlighting to NER and span UI: speedy token-based annotations and precise character highlighting! 🚀 111
Explosion 💥 @explosion.ai · 25/10/2023That means that Prodigy now has 5 official plugins! The Prodigy Docs have also been updated to reflect this change. You can see all the details here: prodi.gy/docs/plugins 110
Explosion 💥 @explosion.ai · 25/10/2023Announcing ✨Prodigy-HF ✨ It's a new plugin that allows you to train @huggingface.bsky.social NER models directly on annotated data in Prodigy. It also provides a recipe to upload annotations to Hugging Face HUB! 254
Explosion 💥 @explosion.ai · 24/10/2023For more details on the Prodigy-PDF plugin and other Prodigy plugins, check our docs page: prodi.gy/docs/plugins 100
Explosion 💥 @explosion.ai · 24/10/2023Want to annotate PDF files with OCR? Our new Prodigy-PDF plugin can help with that! To help explain how to use PDF segmentation and OCR @koaning.bsky.social made a small demo video to highlight the new feature 👀 www.youtube.com/watch?v=rwyz... 161
Explosion 💥 @explosion.ai · 20/10/2023We recently released ✨ Prodigy-ANN ✨ that allows you to use contextual search to find relevant subsets of data to annotate first. To help explain this new feature, @koaning.bsky.social made a small demo to highlight the new feature 👀 youtu.be/jyu2nbjwfXw 121
Explosion 💥 @explosion.ai · 19/10/2023The new Prodigy-ANN recipes can be used to query your collection of images. Once you've created an index you can query it and even directly annotate it with our image.manual recipe. All the details are explained here: 📖 prodi.gy/docs/plugins... 100
Explosion 💥 @explosion.ai · 19/10/2023The new OCR feature uses Pytesseract under the hood to attach parsed text to segments that you've annotated with the `pdf.image.manual` recipe. If you want to learn more, the docs have plenty of extra info: 👀 prodi.gy/docs/plugins... 101
Explosion 💥 @explosion.ai · 19/10/2023✨ Exciting news! We have just released new versions of the Prodigy-ANN and Prodigy-PDF plugins! ✨ With these updates, you can now perform OCR on annotated PDFs and utilize CLIP embeddings to retrieve images that you want to annotate first. prodi.gy/docs/changel... 110
Explosion 💥 @explosion.ai · 18/10/2023As with all Prodigy commands, Prodigy's IAA commands can be run on the terminal using prodigy and the recipe's name. It requires the name of the Prodigy data (e.g., ner) or alternatively can be a flat file like a jsonl file. 110
Explosion 💥 @explosion.ai · 18/10/2023Three types of document-level annotations are supported by the Prodigy IAA commands: 1. Binary: assign a single label 2. Multiclass: assign a single label from a selection of exclusive choices 3. Multilabel: assign one or more labels from a selection of non-exclusive labels 020
Explosion 💥 @explosion.ai · 18/10/2023✨ Context is key in IAA metrics ✨ Consider an imbalanced text classification scenario like detecting fraud. Annotators may show 99% percent agreement, but it's misleading if one always claims the rare class is absent. So raw metrics alone may not tell the full story. 🔍 110
Explosion 💥 @explosion.ai · 18/10/2023Ensuring annotator agreement is crucial in ML training. Disagreements signal issues like ambiguity or insufficient annotator coaching. IAA metrics help fine-tune guidelines early in the project, fostering a shared understanding for accurate dataset annotation. 🔄 120
Explosion 💥 @explosion.ai · 18/10/2023Curious if your annotators are on the same page? Prodigy has just released v1.14.3 with built-in inter-annotator agreement (IAA) metrics to track and measure their agreement. In this 🧵, we'll review Prodigy's document-level IAA metrics. prodi.gy/docs/metrics 153