What did Sentence-Bert return here?

What did Sentence-Bert return here?

Manage alerts

Loading saved threads...

user900476 · External communityPost link
External question — Data Science Stack Exchange Author: user900476 Original post: https://datascience.stackexchange.com/questions/112402 License: CC BY-SA 4.0 — https://creativecommons.org/licenses/by-sa/4.0/ Adaptation: HTML converted to plain text; contact email addresses removed. I used sentence bert to embed sentences from this tutorial https://www.sbert.net/docs/pretrained_models.html from sentence_transformers import SentenceTransformer, util model = SentenceTransformer('all-mpnet-base-v2') This is the event triples t I forgot to concat into sentences, [('U.S. stock index futures', 'points to', 'start'), ('U.S. stock index futures', 'points to', 'higher start')] model.encode(t) returns a 2d array of shape (2,768), with two idential 768-dimension vectors, and its value is different from both model.encode('U.S. stock index futures') and model.encode('U.S. stock index futures points to start') . What could possibly have it returned? It is the same situation for other models on huggingface such as https://huggingface.co/sentence-transformers/stsb-distilbert-base
Quote
Report
Pushpam Punjabi · External communityPost link
External answer — Data Science Stack Exchange Author: Pushpam Punjabi Original post: https://datascience.stackexchange.com/a/112405 License: CC BY-SA 4.0 — https://creativecommons.org/licenses/by-sa/4.0/ Adaptation: HTML converted to plain text; contact email addresses removed. There are two valid inputs to MPNet's tokenizer: Union[TextInputSequence, Tuple[InputSequence, InputSequence]] When you give a list of tuples as an input, from each tuple only the first two sentences, i.e. "U.S. stock index futures" and "points to" are used to encode, similar to BERT's next sentence prediction pre-training task. This is tokenized and converted to input_ids with special tokens as follows: ['<s>', 'u', '.', 's', '.', 'stock', 'index', 'futures', '</s>', '</s>', 'points', 'to', '</s>'] <s> indicates start of sentence and </s> indicates end of sentence. I'm not sure about the </s> before 'points', but this is the output from the model's tokenizer. As such NLU models can work with one or two sentences together, these special tokens are added to identify how many sentences are in the model and they help to separate the sentences. Thus, you get two same vectors from model.encode(t) If you want the same vector without using a tuple, the encode function should contain the following input string: "U.S. stock index futures </s> </s> points to"
Quote
Report

Post Reply

Quoted from Forex.com.bd-Editorial External answer — Data Science Stack Exchange Author: Pushpam Punjabi Source score (net votes, not local likes): 2 Original post: https://datascience.stackexchange.com/a/112405 License: CC BY-SA 4.0 — https://creativecommons.org/licenses/by-sa/4.0/ Adaptation: HTML converted to plain text; contact email addresses removed. There are two valid inputs to MPNet's tokenizer: Union[TextInputSequence, Tuple[InputSequence, InputSequence]] When you give a list of tuples as an input, from each tuple only the first two sentences, i.e. "U.S. stock index futures" and "points to" are used to encode, similar to BERT's next sentence prediction pre-training task. This is tokenized and converted to input_ids with special tokens as follows: ['<s>', 'u', '.', 's', '.', 'stock', 'index', 'futures', '</s>', '</s>', 'points', 'to', '</s>'] <s> indicates start of sentence and </s> indicates end of sentence. I'm not sure about the </s> before 'points', but this is the output from the model's tokenizer. As such NLU models can work with one or two sentences together, these special tokens are added to identify how many sentences are in the model and they help to separate the sentences. Thus, you get two same vectors from model.encode(t) If you want the same vector without using a tuple, the encode function should contain the following input string: "U.S. stock index futures </s> </s> points to"

Cancel quote

Checking account access…