14
I'd like to compute the similarity between two lists of various lengths.
eg:
listA = ['apple', 'orange', 'apple', 'apple', 'banana', 'orange'] # (length = 6)
listB = ['apple', 'orange', 'grapefruit', 'apple'] # (length = 4)
as you can see, a single item can appear multiple times in a list, and the lengths are of different sizes.
I've already thought of comparing the frequencies of each item, but that does not encompass the size of each list (a list that is simply twice another list should be similar, but not perfectly similar)
eg2:
listA = ['apple', 'apple', 'orange', 'orange']
listB = ['apple', 'orange']
similarity(listA, listB) # should NOT equal 1
So I basically want to encompass the size of the lists, and the distribution of items in the list.
Any ideas?