This week, we took a more critical look at the process of discovery in collections and what methods are being tested to make it easier to find things in databases of visual information. It was interesting to look at collection access from a user experience design point of view; facilitating use of the materials is what I believe the core function of any collection should be, but often it is swept to the side with the assumption that anyone who gets as far as your search bar already knows exactly what they are looking for. This problem is often further exacerbated by the complexities of searching for images, whose indexing relies almost exclusively on the information found in their metadata. For example, if you were trying to trace the origin of the pigeon-breast silhouette of early 1900s women’s clothing, there is a slim chance that searching “pigeon breast” would bring up every example of this silhouette in a textile collection. In order for a catalog entry to show up with that search term, its cataloger would have needed to add those words to its description in some shape or form via tags, subject headings, or plain-text description. This is where semantic searching comes in, which requires some level of cultural understanding of a query on the part of an information retrieval system and not just simple 1:1 text matching. Unfortunately, while many commercial search engines utilize this function, not many search algorithms in collections management software are advanced enough to do so.
We looked at two projects in particular that took a machine learning/neural network approach to ameliorating some of these difficulties: “Training the Archive” and IMGS.AI. Both projects center on the central idea of clustering as a form of determining relevance, which they are able to do via image recognition tools that can pull out details such as color, shape, and other purely visual features. To be quite frank, I was not very impressed with the use cases of either of these methodologies. Both projects identify the need for a method of determining visual similarity in search that does not necessarily rely on metadata, and I do agree that this is especially true for large volumes of images, but in my opinion the question of “what can we do with AI tools” seems to have taken precedent over “what problems are people actually having while navigating databases.” IMGS.AI uses sound information retrieval theory with a clever way of instantaneously incorporating user evaluation, but in practice, its interface is highly unclear and essentially amounts to a “hot or cold” game with the database. “Training the Archive” uses a similar clustering system to replicate the knowledge and aesthetic sensibilities of a curator, and while it is true that it was able to pick out some stylistically similar images, I think it’s a bit bold to jump to the claim that it has the potential to become “The Unbiased Curator’s Machine.” It’s an interesting thought experiment, but a very loaded way to present a tool that has its own set of problems with real-world consequences.
The project I thought was most successful was Mitchell Whitelaw’s “generous interface” for the Australian Prints and Printmaking database. With options for guided browsing and search, there are many different entry points into the same collection. In particular, I thought the subjects explorer was a fantastic visualizer for similarity that made image metadata abundantly clear in its interface. The more related a topic was, the more saturated it appears in the sidebar menu, making it very easy to jump between subjects and filter down results. While it requires structured inputs and may not be as flexible as the previous examples, I think the ultimate result is much more navigable and offers more utility to the way people are actually searching and browsing for materials.

I wish more repositories had a way to visualize links between people, organizations, topics, and collections that was so simple to use! The downside to this approach though, as with many tools that yield exciting results, is that it is highly labor intensive and that labor is practically invisible. For each of the 40,000+ works in this collection, each one already had clear, structured metadata with those links between artists and works already established. This is fantastic for a small to mid-sized collection such as this, but this process must be nearly impossible to scale (in a timely manner) for large institutions. For example, the collections I work with at the City of Raleigh currently number at over 35,000+ items accumulated since 1992. I can say from experience that I am constantly updating old descriptions, adding tags and subject headings, creating new scans and photos, and measuring objects in the collection, and there are still at least hundreds of subpar catalog entries that would lead researchers directly into dead ends. And that is just the size of a small local history museum- the North Carolina Museum of History has over 150,000 items in its collection, and huge institutions like the Smithsonian have millions. The sheer amount of labor required to standardize that volume of data is truly astronomical, and I’m not sure how we can create that level of utility and simplicity without either massive funding or the exploitation of GLAM workers- either directly or by training new tools on their work.
I’ll admit, when searching for more collections interfaces to explore, I had a difficult time finding one to share! Not many had very robust exploration options, and most limited browsing to a collection-by-collection basis. I ended up looking at the collections website for the Indianapolis Museum of Art (my hometown art museum as a child), and found it to be pretty intuitive for browsing. On the landing page, there’s an assortment of featured works along with links for popular searches. I hadn’t seen a museum site that had that feature or that exact wording before, but it shows that they have a knowledge of what bridges already exist between their audience and their collection and a desire to encourage those bridges.

There are also easily accessible sections that show only open access or public domain images. The catalog is extremely clear about copyright, which can be so helpful. Often there is just a vague “the creator owns this” statement on catalog entries that is baked into cataloging systems.


I also appreciated that you can view past exhibitions from an exhaustive list, and that all works from the exhibition are represented. Most recent exhibition collections also have images of all of the works used where copyright allows, and catalog entries for every image contain information about its copyright. I wouldn’t say it’s a perfect interface, but it does have some pretty solid methods of exploration.

Bönisch, Dominik. “The Curator’s Machine: Clustering of Museum Collection Data through Annotation of Hidden Connection Patterns between Artworks.” International Journal for Digital Art History, no. 5 (2020): 5.20-5.35. https://doi.org/10.11588/dah.2020.5.75953.
Offert, Fabian, and Peter Bell. “IMGS.AI.: A Multimodal Search Engine for Digital Art History.” International Journal for Digital Art History, no. 9 (2023): 5.28-5.39. https://doi.org/10.11588/dahj.2023.9.91295.
Whitelaw, Mitchell. “Generous Interfaces for Digital Cultural Collections.” Digital Humanities Quarterly 009, no. 1 (2015).



