LISTSERV mailing list manager LISTSERV 16.5

Help for CODE4LIB Archives


CODE4LIB Archives

CODE4LIB Archives


CODE4LIB@LISTS.CLIR.ORG


View:

Message:

[

First

|

Previous

|

Next

|

Last

]

By Topic:

[

First

|

Previous

|

Next

|

Last

]

By Author:

[

First

|

Previous

|

Next

|

Last

]

Font:

Proportional Font

LISTSERV Archives

LISTSERV Archives

CODE4LIB Home

CODE4LIB Home

CODE4LIB  July 2026

CODE4LIB July 2026

Subject:

Re: llm technology is especially useful

From:

Tim Spalding <[log in to unmask]>

Reply-To:

Code for Libraries <[log in to unmask]>

Date:

Thu, 9 Jul 2026 11:54:28 -0400

Content-Type:

text/plain

Parts/Attachments:

Parts/Attachments

text/plain (262 lines)

There are many conversations to be had about the value and dangers of AI. I
don't discount anyone's opinion. My own is AI has many good and important
uses, many bad—and is a particular disaster for education.

But Code4Lib is at a minimum about code, and AI coding has absolutely
transformed the software development industry over the last six months. If
it hasn't hit you yet, it will. It is the most significant change in
decades. It would be short sighted to shut down or stigmatize the
discussion of it here.

Tim Spalding
Founder, LibraryThing


On Thu, Jul 9, 2026 at 11:39 AM Karen Coyle <[log in to unmask]> wrote:

> This reminds me of what we were told in school when calculators first
> became affordable:
>
> "You need to study math and learn to use it even though you now have a
> calculator. Only if you understand math will you be aware if some
> mistake has been made."
>
> Like when I hit the multiply key rather than the divide key.
>
> I'm good with using computers, whether than LLM or some other software,
> for crunching data. I'm much less confident that I can trust computers
> to operate with human language. Crunched data can be checked, and the
> output is factual. Language is not. I'm back to "the right tool for the
> job" but also "understand the limits of your tools."
>
> kc
>
> On 7/8/26 2:53 PM, Vishal Patel wrote:
> > Given Alex’s strong response, I feel compelled to weigh in.
> >
> >    1.
> > It’s clear that LLMs are divisive. We (namely, Mackenzie Salisbury)
> started one dedicated to LLMs at [log in to unmask]
> - it leans more LLM-curious.
> >    2.
> > I don’t detect any “LLM slop” in this thread. This appears to me be a
> constructive evaluation of an LLM case study.
> >    3.
> > I would advise against snubbing or canceling a forum or a person if you
> (even by mistake) catch a whiff of LLM generated content.
> >       *
> > Language, ultimately, is just a collection of symbols to which we have
> assigned meaning, and, when we find ourselves have strong emotional
> reactions to those symbols, the empirical evidence indicates that it will
> serve us - individually and collectively - better if we learn to be curious
> about those emotions, rather than checking out or leaving the group
> entirely.
> >       *
> > It is true that today’s symbols are rapidly commingling with
> machine-generated symbols (linguistically, visually, & audibly), but
> librarians have always existed to serve as stewards of collections of
> symbols for the community. If we opt out when we see or hear a new dialect
> (e.g. AI-infused English) being written or spoken, not only is that
> antithetical to the professional values of librarianship, but it fractures
> our professional body.
> >
> > Going back to the purpose of Eric’s thought starter, I am greatly
> appreciating the value of LLMs. As an example, within the paste 24 hrs, I
> was able to:
> >
> >    1.
> > Write SQL queries in BigQuery to analyze 6 months of traffic & search
> patterns on Lane library’s website
> >    2.
> > Analyze 6 years of historical data from our LibAnswers & LibCal systems
> >    3.
> > Analyze 16 years worth of proxy server data
> >    4.
> > Synthesize all analyzes into a strategic plan
> >    5.
> > Summarize the findings as an infographic to share with my executive team
> >
> > As a former data scientist, I can attest that this type of analysis
> would have previously taken me 2-3 weeks to assemble, and the polished data
> visualization would have taken another week itself. LLMs make easy work of
> data analysis & coding.
> >
> > Sincerely,
> >
> > Vishal Patel, MD, PhD
> > Director, Digital Strategy & Library Technology
> >
> > Lane Medical Library
> >
> > Stanford School of Medicine
> >
> >
> > 300 Pasteur Drive, L109, Stanford, CA 94305-5123
> >
> > [log in to unmask]<mailto:[log in to unmask]>
> >
> > (650) 723-7196 *41425
> >
> > Pronouns:  he, him, his
> >
> > [cid:[log in to unmask]]
> >
> > From: Code for Libraries <[log in to unmask]> on behalf of Alex
> Dunn <[log in to unmask]>
> > Date: Wednesday, July 8, 2026 at 2:18 PM
> > To: [log in to unmask] <[log in to unmask]>
> > Subject: Re: [CODE4LIB] llm technology is especially useful
> >
> > I think it's high time for this mailing list to enact a ban on LLM
> > slop.  In the meantime, I'm unsubscribing.  See you all around.
> >
> > On Wed, Jul 8, 2026 at 8:39 AM Eric Lease Morgan
> > <[log in to unmask]> wrote:
> >> On Jul 4, 2026, at 7:44 PM, Karen Coyle <[log in to unmask]> wrote:
> >>
> >>> I would feel better about this if these results didn't sound like the
> platitudes of marketing speak. ("Collaborative refinement of library
> services" is something I don't think any of us would say about C4L.) Where
> did the model get this tripe?
> >>
> >> TL;DNR - I advocate the use of LLM technology in libraries. It can be a
> supplement to our existing processes, not a replacement.
> >>
> >>
> >> Thank you for the reply because I really feel our community good
> benefit from discussion on these topics.
> >>
> >> Tripe? I had to look up the definition of that word. Where did the
> result get such a response? Many places, but one of the more significant is
> my locally configured "system prompt". What's that? A system prompt is an
> extra little bit, behind the scenes configuration sent to a large-language
> model (LLM). My current system prompt follows:
> >>
> >>    Return results as if they were written by a student
> >>    attending a liberal arts college. Ask questions, sometimes,
> >>    but not always. The model is working within a generative-AI
> >>    system called a RAG, and therefore results are intended to
> >>    be primarily drawn from the underlying MCP system; results
> >>    drawn from outside the system are to be kept to a bare
> >>    minimum. The model is intended to be used as analysis tool
> >>    not an oracle. Do not voice results in the first person!!!
> >>    When citing sentences, include item and index values.
> >>
> >> If my system prompt said something like "Voice replies as if written by
> an eighth grader", then the results would be expressed differently.
> Differences in system prompts make a significant difference in results. You
> should see the sort of things I get back when I specify second graders or
> erudite college professors.  :-D
> >>
> >>
> >>> But what really concerns me is that the system is a black box (others
> have noted this), which means that there is no way to evaluate the result
> other than ones' gut feeling that it's "right." Your gut feeling returns
> "true and accurate" while mine concludes that honestly some of the "facts"
> in the statements below could be wrong. For all that C4L has great
> discussions, I don't see "code sharing" as taking place often in the body
> of the emails. (Maybe a snippet or two, but not as stated here.) The last
> paragraph on MARC doesn't convince me much. The phrase "annoying data
> format" comes up only once when I search on the C4L archive, and it's in
> one of your posts, Eric. I wouldn't include this in a summary of C4L list
> users' statements on MARC even though we are pretty critical of it. I also
> have doubts that one can conclude from the list that the group is "focusing
> heavily on MARC records—the standard format for library catalog data."
> There is a fair amount of discussion about MARC but have posters actually
> said here that it's the standard format for library catalog data? (We
> probably assume that everyone here knows that.) Could the software have
> gotten that from elsewhere, given that there is a fair amount of
> documentation online?
> >>>
> >>> --
> >>> Karen Coyle
> >>> [log in to unmask]
> https://urldefense.com/v3/__http://kcoyle.net__;!!G92We9drHetJ8EofZw!emYRrTgEe2gQVprd47jFZCCIprO0hxT1USRlq7XMgipz5vbj2Jf6kauD7i8DN_ZAehOjotejg5xnn7JIIk41Ww$
> >>
> >> Granted, above is more difficult; I more or less (mostly more) agree,
> but...
> >>
> >> The box is more gray than black. The whole process is rooted in the
> computation of geometric distances between words mapped in a VERY large
> n-dimentional space. To elaborate, it works something like this. A HUGE
> pile of text is accumulated. The text is parsed into tokens (think "words")
> to create a vocabulary. This results in a matrix where each row (millions
> or billions of them) is a document, and each column is a token from the
> vocabulary. At the intersection of each row & column is a measurement, and
> the measurement might be the number of times the given token is found in
> the given row. This results an a GREAT BIG set of vectors "pointing" to
> locations in the space. Given some input (at least a word or better yet a
> large set of words), the input is vectorized in the same manner as the
> original LLM. The vector is compared to all of the other vectors in the set
> to identify most similar vectors, where similarity is denoted to something
> like cosine distance. This is a process of linear algebra, and it is a kind
> of find or search process. Works in the manner similar to auto-completion
> or auto-correct but on a REALLY big scale.
> >>
> >> For example, in English, the word "the" is very frequently followed by
> an adjective but ultimately by a noun. These patterns are manifested in the
> LLM.
> >>
> >> Using the sort of process outlined above, given words are mapped with
> similar words, and the similar word are output. It does not necessarily
> identify nor extract exact phrases from text unless specifically asked to
> do so. Also, in my example, I only read one month's of Code4Lib archives,
> and in that month there may have been more discussion about MARC than not.
> In any event, the process does work (for the most part), and it is a
> real-world application of something first articulated by a man named John
> Firth around 1957 who said, "You shall know a word by the company it keeps"
> -- context.
> >>
> >> It is a black box? Yes, mostly. It black in the same way Google's
> search algorithms are black. It is back in the same way our bibliographic
> indexes rank relevance. It is as black as relational database
> implementations perform join queries through many-to-many relationships.
> >>
> >> Very important: I do not advocate the use of LLMs sans context. In
> other words, I do not advocate asking very general questions, like "What is
> the best Shakespeare play?" to LLMs because the only context included in
> the response is what is in the model. On the other hand, I very much
> advocate the use of retrieval-augmented generation (RAG) and/or model
> context protocol (MCP) servers. These tools get input from known and
> verifiable collections of text. The results of RAG and/or MCP queries are
> sent to an LLM for interpretation, and well-implemented results will
> include ways to backtrack results to the source. I think the use of RAG
> and/or MCP technologies applied to library collection can be very useful.
> LLM are not panaceas though.
> >>
> >> For example, the other day I queried various bibliographic indexes for
> the phrase 'big science'. This resulted in 1,200 abstracts from scholarly
> journals for a total of 350,000 words (which is bigger than Moby Dick). I
> then RAG-ed against it to extract definitions of 'big science', learned who
> helped define it, and how that definition changed over time. It worked. It
> worked well. It saved me a whole lot of time. It supplemented my reading
> process, not replaced it.
> >>
> >> I have learned two additional things. First and foremost, I take the
> results as plausible, not truth. Thus, I like to believe I practice
> information literacy along the way. Second, I am always dubious of the
> adjectives returned by LLM, especially the superlatives.
> >>
> >> What I would really like to see is the creation of one or more LLM
> built by the library profession. This way we would know whence the model
> came, how it was implemented, and remove the blackness. Such an effort
> would be akin to collaboration we have seen in the past when it comes to
> collection building or metadata sharing. Yes, it would be very expensive,
> but if we were to pool our resources, then I think it could be done.
> >>
> >> Lastly, those were a lot of words. Thank you for listening. I hope the
> discussion continues.
> >>
> >> --
> >> Eric Lease Morgan, Librarian Emeritus
> >> University of Notre Dame
>
> --
> Karen Coyle
> [log in to unmask]
> http://kcoyle.net
>
>

-- 
Check out my library at https://www.librarything.com/profile/timspalding

Top of Message | Previous Page | Permalink

Advanced Options


Options

Log In

Log In

Get Password

Get Password


Search Archives

Search Archives


Subscribe or Unsubscribe

Subscribe or Unsubscribe


Archives

October 2026
September 2026
August 2026
July 2026
June 2026
May 2026
April 2026
March 2026
February 2026
January 2026
December 2025
November 2025
October 2025
September 2025
August 2025
July 2025
June 2025
May 2025
April 2025
March 2025
February 2025
January 2025
December 2024
November 2024
October 2024
September 2024
August 2024
July 2024
June 2024
May 2024
April 2024
March 2024
February 2024
January 2024
December 2023
November 2023
October 2023
September 2023
August 2023
July 2023
June 2023
May 2023
April 2023
March 2023
February 2023
January 2023
December 2022
November 2022
October 2022
September 2022
August 2022
July 2022
June 2022
May 2022
April 2022
March 2022
February 2022
January 2022
December 2021
November 2021
October 2021
September 2021
August 2021
July 2021
June 2021
May 2021
April 2021
March 2021
February 2021
January 2021
December 2020
November 2020
October 2020
September 2020
August 2020
July 2020
June 2020
May 2020
April 2020
March 2020
February 2020
January 2020
December 2019
November 2019
October 2019
September 2019
August 2019
July 2019
June 2019
May 2019
April 2019
March 2019
February 2019
January 2019
December 2018
November 2018
October 2018
September 2018
August 2018
July 2018
June 2018
May 2018
April 2018
March 2018
February 2018
January 2018
December 2017
November 2017
October 2017
September 2017
August 2017
July 2017
June 2017
May 2017
April 2017
March 2017
February 2017
January 2017
December 2016
November 2016
October 2016
September 2016
August 2016
July 2016
June 2016
May 2016
April 2016
March 2016
February 2016
January 2016
December 2015
November 2015
October 2015
September 2015
August 2015
July 2015
June 2015
May 2015
April 2015
March 2015
February 2015
January 2015
December 2014
November 2014
October 2014
September 2014
August 2014
July 2014
June 2014
May 2014
April 2014
March 2014
February 2014
January 2014
December 2013
November 2013
October 2013
September 2013
August 2013
July 2013
June 2013
May 2013
April 2013
March 2013
February 2013
January 2013
December 2012
November 2012
October 2012
September 2012
August 2012
July 2012
June 2012
May 2012
April 2012
March 2012
February 2012
January 2012
December 2011
November 2011
October 2011
September 2011
August 2011
July 2011
June 2011
May 2011
April 2011
March 2011
February 2011
January 2011
December 2010
November 2010
October 2010
September 2010
August 2010
July 2010
June 2010
May 2010
April 2010
March 2010
February 2010
January 2010
December 2009
November 2009
October 2009
September 2009
August 2009
July 2009
June 2009
May 2009
April 2009
March 2009
February 2009
January 2009
December 2008
November 2008
October 2008
September 2008
August 2008
July 2008
June 2008
May 2008
April 2008
March 2008
February 2008
January 2008
December 2007
November 2007
October 2007
September 2007
August 2007
July 2007
June 2007
May 2007
April 2007
March 2007
February 2007
January 2007
December 2006
November 2006
October 2006
September 2006
August 2006
July 2006
June 2006
May 2006
April 2006
March 2006
February 2006
January 2006
December 2005
November 2005
October 2005
September 2005
August 2005
July 2005
June 2005
May 2005
April 2005
March 2005
February 2005
January 2005
December 2004
November 2004
October 2004
September 2004
August 2004
July 2004
June 2004
May 2004
April 2004
March 2004
February 2004
January 2004
December 2003
November 2003

ATOM RSS1 RSS2



LISTS.CLIR.ORG

CataList Email List Search Powered by the LISTSERV Email List Manager