How can I perform entity extraction in Pipeline Builder ? If I have documents, or transcripts of documents, what prompts should I use and how to structure the output to extract specific pieces of information, if present ?
When using an LLM node in Pipeline Builder to perform Entity extraction, it is possible to use a prompt that asks the LLM to return key-value pairs, which allow to easily evolve entities to extract.
Example of prompt for extracting entities in Pipeline Builder
A prompt like the following.
You are a tool for <industry> extract information from <type of documents>
You are looking for <type of information>, these are <description>
Your goal is to parse the following page text into an array of JSON structs
where every entry contains the following fields:
- <key>: <description of what to extract, format, type>
- <key>: <description of what to extract, format, type>
- <key>: <description of what to extract, format, type>
If the text doesn't contain an a <type of information> return null
The output should be a valid JSON (without \n or or other escaping)
Tip: You can as well add the document itself (for vision models) directly by adding the media reference in the prompt !
You can then extract those key/values via the below transforms

