Wednesday, February 23, 2011

Test Bed for (X)HTML Conventions for Scholarly Publication

The main reason I joined the Institute for the Study of the Ancient World at NYU was to be part of initiating a program of digital publication of peer-reviewed scholarship. We haven't announced anything formally and this blog post isn't that announcement. It is the beginning of a nuts-and-bolts conversation about the markup of digital scholarship that is intended to encourage long-term viability, flexible re-use, and easy display (among many other things).

To get right down to business, http://dl.dropbox.com/u/17002562/isaw-papers-preprint.xhtml is the very temporary URL for a preprint version of "Review of Ptolemaic Numismatics, 1996 to 2007" by Catherine Lorber and Andrew Meadows. I'm very grateful to Andy and Cathy for their willingness to be part of this experiment. Their work is largely done. Now it's up to me to make progress on the markup and I'm hoping to do that in a very public way.

But where to begin the conversation? I think the best approach is to admit I'm in the middle of things and just start laying out issues and thoughts. Keep in mind that everything is subject to change...
  • The format for ISAW digital publications is XHTML with RDFa. XHTML (for now 1.1 but moving to XHTML5) is a widely supported standard with excellent tooling that is directly viewable in many contexts. That makes it appropriate for long-term archival storage of born-digital scholarship.
  • Internal reference structures are important.For now this means each <p> element has an @id. div's of class 'section' also have @id attributes. This is in anticipation of using the semantic elements of HTML5.
  • Named entities will be tagged with links to stable resources describing those entities. For geography, Pleiades. For many other entities, Wikipedia. See below for RDFa patterns.
  • Existing ontologies/vocabularies will be used whenever possible. Geographic entities are typed as "dcterms:Location". That sort of thing.
  • Basic constructs for marking up bibliography and footnote-like structures are lacking for HTML-based markup languages. There are lots of semi-complete "best practices" but narrowing these down to a consistent and flexible convention will be an importnat process.

Looking ahead:
  • Multiple formats will be supported. We will distribute this text as "raw" valid xhtml. It will be hosted in a more interactive environment that does slick things like make maps, etc. Epub, pdf... all those are coming. Again, the ease with which a base XHTML representation can be converted to these other formats is one reason to use XHTML.
  • We'll use CC licenses Right now the document is CC-BY-NC-ND. We'll drop the ND eventually, perhaps the NC as well. The preprint is ND as a signal that a better version is coming from us.


A word on RDFa (a standardized way of embedding information in XHTML pages)...

The basic pattern that I'm using to markup named entities is illustrated by the sentence:
In a study of tax receipts from early Ptolemaic <a class="citation"
href="http://pleiades.stoa.org/places/991398"
typeof="dcterms:Location" rel="iana:describedby"
property="rdfs:label">Thebes</a>...


That produces the RDF/Turtle
[ a dcterms:Location ;
rdfs:label "Thebes"@en ;
iana:describedby <http://pleiades.stoa.org/places/991398>].
You can see the turtle for the whole document at http://bit.ly/hJjgcx.

An "English" equivalent of the turtle snippet is 'There is a site in the text with label "Thebes" and a description at http://pleiades.stoa.org/places/991398.'

I like the use of the 'describedby' @rel value here. It's defined in the IANA's register of rel values (http://www.iana.org/assignments/link-relations/link-relations.xml). I take the semantics to be "I'm not saying I'm linking to Thebes itself, only to a description of it." That seems nice and "semantic webby".

There's more to come but I'm getting this out there just to get the ball rolling...

Monday, February 21, 2011

Quick poll: Worldcat, Library of Congress, or Both

There are lots of ways of encoding bibliographic data on the web, but this post isn't about that problem. Instead, I'm wondering what is "the community's" preference between Worldcat and the Library of Congress when creating Semantic Web/Linked Open Data references.

As an example, the URIs http://www.worldcat.org/oclc/829279 and http://lccn.loc.gov/74155758 each lead to information about John Hayes' Late Roman Pottery published in 1972.

Which one of these is preferable as the long-term description of this volume? Worldcat or LOC. The use-case is a digital publication with bibliography that ideally includes a link to one or the other or both for all printed volumes or other appropriate entities.

Perhaps a discussion will ensue in the comments but here are some quick issues:
  • There are multiple URIs for that one volume in Worldcat. http://www.worldcat.org/oclc/462730938 gets you to the Danish Union Catalog.
  • There are still concerns about the licensing of Worldcat data.
  • The LOC record is to a physical volume in a single national library and may not be intended as a description of the abstract concept (e.g. a FRBR Work). I don't know that Worldcat URIs solve this problem but they have the implication of a higher level of abstraction.


Votes and/or comments are appreciated.

Monday, February 7, 2011

Quick poll: Wikipedia or DBPedia?

I've created a poll near the upper right of this page. In longer form: when making persistent "Linked Data/Semantic Web" references to concepts described in Wikipedia, is it "best practice" to link to Wikipedia or to DBPedia? As in, "http://en.wikipedia.org/wiki/Augustus" or "http://dbpedia.org/resource/Augustus"?

Friday, February 4, 2011

Access to Roman Art: Observations by Peter Stewart

The last few times I've gone to speak about issues of scholarly communication/digital humanities/digital archaeology/etc, I've opened up with a quote from Peter Stewart's 2008 book The Social History of Roman Art [Worldcat]. That's a great little book, and I was particularly pleased when reading it that Stewart is explicit about the effects of access to evidence and images on his selection and narrative. And I was further pleased that he talks about his personal efforts to solve those problems. I'll illustrate this by a series of passages given in their order of appearance:
Unfortunately, my comments in the Introduction about the problems of acquiring images were born out in the book's preparation, and I had very considerable difficulties and delays in acquiring most of the images reproduced here. I therefore owe a special debt to those who helped me to obtain pictures, and to those image-providers who waived or reduced reproduction fees. (p. xv)
Then from that introduction:

To an extent, however, these are all obvious problems of evidence and interpretation which are familiar in any branch of historical study. Other problems are insidious and lie unremarked in the methodological hinterland of books like this one. I have said that the use of examples must be highly selective. But behind any book on Roman art, there are processes of selection that are largely beyond the author’s control. Most Roman art historians will never, in their lifetime, see more than a tiny percentage even of the more significant works that survive. This is not simply because of the magnitude of this great body of material. It is also because most pieces are inaccessible. Many of the finest and most interesting Roman antiquities are in private collections, and many of these are unpublished, sometimes because of scholars’ anxieties about the legality of their origins. However works preserved in museums can be at least as difficult to access. Few museums are able to exhibit more than a small minority of the objects they hold. It is not infrequent (or surprising) for some of the objects in storage to be, effectively, lost, and for other reasons it may be hard for specialists to see material, particularly if it has been excavated recently. New discoveries may take many years to become familiar within the field, and even longer to filter into general, synoptic studies of Roman art.

So, for a variety of reason, authors depend heavily on other people's publications of Roman art, where they exist, and on their illustrations. The photographs themselves are usually supplied by the museums that own the work concerned, or simetimes by commercial agencies. In many cases no photograph exists, and new photography may not be permitted. In other cases, the acquisition of photographs proves lengthy or impossible. Moreover, the photographs (especially colour images) and the permission to reproduce them in print can be extremely costly both for individual authors and for their publishers. (p. 8)

The passages need to be read in context. It's not an angry book, and these introductory are comments are followed by interesting and challenging extended essay on the topic indicated by the title. I can highly recommend it. But back to the issue of access, here's a passage from the ending Bibliographical essay:
Finally, the photo-sharing website flickr.com contains thousands of images relevant to Roman art, many of them with 'Creative Commons' copyright licenses that make them easy to use legitimately for, e.g. educational purposes. Within that site the 'Chiron' group especially is dedicated to making images available for classical teaching and research. This site carries many of my own photographs (under the screen name 'Tintern'), including colour images of the House of the Vettii and other sites mentioned in this book. (p. 174)
So mad props to Dr. Stewart for raising the issue of access and then doing something about it. A book from CUP in which the author cites his flickr.com account? That's progress.

Monday, January 31, 2011

In-house commenting systems may not be necessary

Somewhat wishy-washy title, I know.

But here's my point, I look forward to a world of stable URIs for intellectual content in which responses to scholarship and primary data are distributed around the Net.

A case in point, my NYU colleague Chuck Jones blogged about the digitization of some of Blegen's diaries by the American School of Classical Studies.

If you look at the bottom of the post, you'll see that he included the Pleiades URI's for both Mycenae and Tiryns.

It is now the case that a Google search for the Tiryns URI lists Chuck's AWOL post.

Assuming that ASCSA doesn't move that resource to a different URI and that the post remains available, stable URIs for Tiryns and Mycenae have now been permanently associated with the ASCSA resource. And that with the publisher of the information doing nothing. (Though it would be nice if ASCSA ditched the "index.php" from their URIs. See here.)

And note that I'm walking a fine line in this post. The Pleiades URIs that Chuck included explicitly in his post don't appear in the text of mine. I don't see any reason to clog up the Google search with this meta-meta-commentary.

By way of slightly living up to the title, my point is that such a decentralized "commenting system" should be encouraged. If you're able to link from your content to a stable URI that more-or-less represents the same concept, do so. And use such URIs when you're talking about other's people's work. That will encourage a distributed network of publication and response that is robust, open and encompasses many forms of expression from tweets, to blogposts, to more formal work, and beyond.

Wednesday, November 3, 2010

Responses to "Progress on Museum URIs"

Three people responded to yesterday's post on museum URIs.

Leif Isaksen left a comment to the effect that he's not too concerned about differing base URIs for museum collections. I agree that there are worse things than the string "collection." in "http://collection.britishmuseum.org/object/YCA62958". The original explanation was to reduce load on an individual server. Without meaning to get too technical, the "/object" can be an effective load reducer by passing requests to a proxy. Bottom line: in an ideal world, I'd drop the "collection.", but I'm not too worked up about it.

Eric Kansa responded on his blog. His point had an interesting overlap with an e-mail I received. I won't quote that in its entirety as the author could have made it public if s/he wanted to. Here's a snippet:
but to me it seems a very bad idea to think that only museums can claim the right to designate URIs for their objects; there should be a standard that can be used by museums as well as by scientists outside of museums...
I took this as responding largely to
2. In order to avoid that everybody invents a new URI for the same
object, there should be one authority known to the whole world that
assigns such a URI.

3. This authority is naturally the museum that keeps the object,
because it is the only institution that can verify that two
different use cases of museum object URIs actually describe the same
thing.
Taking Eric's and Anonymous' comments together, I read them as calling for a multi-vocal internet in which many agents can assert an identity for an object, with those identities together forming a distributed and diverse commentary on the human past. I totally agree. To be self-critical, I may well have mis-read M. Doerr's e-mail. If he's calling for recognition of the exclusive right of museums to identify their objects, that's a non-starter. It's neither the right thing to do nor is it possible. On first reading, I took his e-mail to represent a welcome assumption of responsibility by museums to provide a locus of stability for reference to their collections. But to be clear, objects will have multiple identifiers. Referring back to a common identifier promoted by and discoverable at the holding institution will ease the process of recognizing that two or more identifiers refer to the "same thing". That will itself promote the idea of a discoverable and multi-vocal discussion about the past.

Tuesday, November 2, 2010

Progress on Museum URIs

I'm including the full text of an e-mail sent by Martin Doerr of the Center for Cultural Informatics on Crete. It's been forwarded to me by a couple of people and there's a call for comment towards the end so it seems to be a public document. That's good because it's an excellent step forward in promoting stable URI's for museum collections. From my perspective, it mostly speaks for itself. Section 7 did cause some concern:
...

Under this consideration, Dominic proposes for the British Museum (http://www.britishmuseum.org/), that all objects of the Museum should be identified on the Semantic Web by the following: http://collection.britishmuseum.org/object/ followed by the "PRN number".

For instance, the Rosetta Stone has the PRN number: YCA62958, hence the "official" URI of the Rosetta stone is: http://collection.britishmuseum.org/object/YCA62958 . This URI should never become direct address of a document.
Just to be clear, if a user cuts-and-pastes 'http://collection.britishmuseum.org/object/YCA62958' into an address bar, or a document links directly to that (which I've just done), that should produce a human readable page. I'd like to see that happen without redirection. If you redirect to that same URL with ".html" appended, then authors will cut-and-paste that string into their documents. If a good non-crufty URI exists, that's what should appear in address bars and that's what should stand as the 'permalink'.

More generally, URIs should promote unity and overlap, not division, between the "semantic web" and the "plain-old web" (POW).

Section 7 also endorses URIs that have a different domain name from the institution itself, e.g. the "collection." in front of "britishmuseum.org". I don't like that. The reason given is to avoid the implications of name changes in the future. Ugh. Institutions should formally endorse the URIs they mint and make them as simple and short as possible. This decision should be taken at the highest levels of the institution. In the BM's case, that may mean the 25-member Board of Trustees.


Finally, the excellent and useful Europeana is mentioned. I'll take this opportunity to note that while http://europeana.eu/portal/record/00401/034BEA5CC6F88ADC6E7DCF5D7C5FECEA8FF85528.html works, http://europeana.eu/portal/record/00401/034BEA5CC6F88ADC6E7DCF5D7C5FECEA8FF85528 doesn't. It should.



Dear colleagues,

I'd like inform you about our discussion today with Dominic Oldman,
Deputy Head of Information Systems, British Museum, his team and
representatives of the Research Space project
(http://sites.google.com/site/rspaceproject/the-team):

1. It is necessary that museum objects are uniquely identified by
suitable URIs in Semantic Web applications.

2. In order to avoid that everybody invents a new URI for the same
object, there should be one authority known to the whole world that
assigns such a URI.

3. This authority is naturally the museum that keeps the object,
because it is the only institution that can verify that two
different use cases of museum object URIs actually describe the same
thing.

4. This URI should be derived in a simple way from the inventory
numbers published in exhibition catalogues, on on-line museum
catalogue access or by asking museum staff, to avoid an error-prone
equivalence matching process.

5. This URI should have a form that enables any museum that wishes
to do so to provide a Linked Open Data service resolving to the
description of that object. Note, that this URI must not be the URL
of an existing document about the object, but it must activate a
standard mechanism prescribed by the Linked Open Data Initiative to
redirect to a document saying what the URI means.

6. This museum object URI will continue be useful for communicating
uniquely about the object, even if the museum never will install an
LoD service, or if the way of dealing with LoD resolution requests
will change.

7. The way to create this URI should be the following: The museum
decides a base URL that will be extended by the inventory number of
the object. The base URL could be within the domain name of the main
museum Website, but in order to stay clear of possible name change
of the latter, a new domain name might be advisable. Also, for
larger museums, resolving LoD access requests to object information
may cause some server load, that can more easily be balanced with a
second name.

Under this consideration, Dominic proposes for the British Museum
(http://www.britishmuseum.org/), that all objects of the Museum
should be identified on the Semantic Web by the following:
http://collection.britishmuseum.org/object/ followed by the "PRN
number".

For instance, the Rosetta Stone has the PRN number: YCA62958, hence
the "official" URI of the Rosetta stone is:
http://collection.britishmuseum.org/object/YCA62958 . This URI
should never become direct address of a document.

It would be good, if Europeana experts to comment, if they regard is
an adequate approach for Europeana, and could transfer this message
to other museums and providers to follow this practice.

I intend to present this on the CIDOC Conference in Shanghai. I
would be very glad if I and Dominic could get a response within the
next week, if you endorse the procedure, and if you will support us
to spread the practice.

If you need further clarifications, please let me know as soon as
possible.

Best wishes,

Martin
--

--------------------------------------------------------------
Dr. Martin Doerr | Vox:+30(2810)391625 |
Research Director | Fax:+30(2810)391638 |
| Email: martin@ics.forth.gr |
|
Center for Cultural Informatics |
Information Systems Laboratory |
Institute of Computer Science |
Foundation for Research and Technology - Hellas (FORTH) |
|
Vassilika Vouton,P.O.Box1385,GR71110 Heraklion,Crete,Greece |
|
Web-site: http://www.ics.forth.gr/isl |
--------------------------------------------------------------

Thursday, October 28, 2010

Ancient Mediterranean Objects at the NMHN

Using posterous.com to track URIs. Here's an Ancient Mediterranean object at the National Museum of Natural History.

If you're reading this at http://mediterraneanceramics.blogspot.com/ , that's part of the experiment as well.

Saturday, October 16, 2010

Change Happens (if it can)

As the result of an e-mail exchange with Neel Smith, one of the designers of the Canonical Text Service Protocol, I've come up with the following formulation:
If a character in a URL can change, it will.
I'm not the only person to think this but I just wanted to get that thought out in simple, direct language.

But what do I mean? Take Worldcat URLs such as http://www.worldcat.org/oclc/502674170. That "www." is annoying and should not be part of the URL that Worldcat presents as its permanent identifier for the book. At some point in the future, somebody there will realize this and remove those unnecessary characters. But http://worldcat.org/oclc/502674170? Now you're talking! And look, it already works.

It's true that the "oclc" could be shortened so maybe I need to qualify the formulation, but I'm not going to for the following reason. Changing those characters would risk collision with other identifying schemes that Worldcat supports such as http://worldcat.org/isbn/0754677737 . The 'www.' is unstable because it can be removed without breaking anything.

The simple formulation stands: If a character in a URL can change, it will.

The implication is, "be aggressive about removing all unnecessary characters from your URLs." The following is a horror-show:
http://www.worldcat.org/title/digital-research-in-the-study-of-classical-antiquity/oclc/502674170
It just looks unstable. Leading me to another formulation:
If a URL looks unstable, it is.

Tuesday, September 21, 2010

Discussing Citation by Example

I've started a set of pages at the Digital Classicist Wiki on the topic of Citation in digital scholarship. In progress, under construction, etc., etc., etc.

The goal is to move existing practice towards a broad understanding of how to make citations to such categories of evidence as primary written sources, geographic entities, cataloged objects, and secondary scholarship so that those citations are:
  • Clearly identified in a robust yet rich fashion
  • Recognizable by automatic agents
  • To resources that are stable over the long-term

But I don't think it will be possible to establish and drive adoption of one very detailed standard. Better to have a simple notation - I follow others in suggesting 'class="citation"' for (x)html - that can indicate the presence of more detailed markup. I'm a fan of RDFa so I further discuss that on the page "Citations with added RDFa.

The Digital Classicist community is pretty open and I'm very grateful to G. Bodard (a.k.a palaeofuturist) for saying the equivalent of "Go for it." when I raised the possibility of hosting these materials in his realm.

There's a category for all the pages and I hope that list will grow.

Wednesday, September 8, 2010

References that just work (but I understand it's not that simple...)

Go to Google. Type in "John 20:24", then hit return. You can even click the "I'm Feeling Lucky" button. Or here's a direct link.

Or try the same thing in Bing (which provides results for Yahoo), and Altavista.

As you'll see, all three searches get you to the relevant passage of the Gospel According to John. And if you poke around on the Biblegateway site, you'll see various translations (but where's the Vulgate?).

That's impressive. It indicates that human readable references can be become so stable that automated agents are able to correctly translate them into links to particular chunks of primary text.

Here are some variations on the theme (all in Google):
"jean 20:24" (at google.fr): Not spot on, but pretty close.
"1 John 2:1": That's a reference to the first epistle of John. Entered into the "address bar" in Chrome. Seems to work.
"John 3": Unqualified chapter reference. Good to go.
"ephesians 1:2": Works.
"eph. 1:2": That abbreviation is OK.
"eph 1.2": Things become fuzzier if I don't use the ':' that is conventional in references to Christian scripture.
"Ephesians 2:4-10": Spans work as well, when properly formatted.

Again, I think this is interesting. Taking the New Testament as a corpus of Ancient Mediterranean texts that were written between the mid-first and third (at the latest: the Epistle of James 1 isn't definitively quoted until Origen) centuries AD makes it relevant to the study of the Ancient World as a whole. As a corpus, it's been around for a long time. Athanasius's letter of AD 367 is one conventional date for the determination of what was in, and what was out.

Those comments aside, the point remains that it is possible to automatically reverse engineer the citation scheme of a very stable corpus. I guess one caveat is that I don't absolutely know that Google, Bing, etc. haven't special cased strings that are plausibly references to the NT. Any ideas?

My larger goal is to think about references to so-called "primary texts" that just work. Given the above, my ad hoc, working definition of "primary text" is any text with a sufficiently stable name and citation scheme that search engines can find it. Sure, that's circular and incomplete, but it will do for now.

Let's try some others:
"gilgamesh 3": Muddled.
"gilgamesh tablet 3": Better.
"Iliad 23": Not bad. No Greek.
"Iliad 23.100": Individual line references don't work.
"Homer Iliad 23.100": Not better.
"Quran 32": I see it as the third link.
"hemingway, the old man and the sea": For comparison. Wikipedia is the top page for me; that's not the text itself. And Amazon is up there, as in the work is in copyright so I'd have to pay. Not sure I want to follow the links that say I can download the text for free.

A major distinction between references to NT texts and the second group is the ability of Google to handle full chapter and verse ('n:n') references. That doesn't seem to work for the Iliad. That's worth exploring.

If I go to Perseus and use the search box at the upper right, "homer iliad 23.100" doesn't work directly. Nor does "iliad 23.100". But "Hom. Il. 23.100" does. If I try that string in Google (link), it gets me to the Chicago version of the Perseus texts (via the 2nd ranked link when I tried it.). [I'll take this opportunity to note that the Chicago Perseus is wicked, and that it's likewise wicked cool that Perseus texts are licensed so this redundancy is possible.]

That kind of variation is one of the reasons I parenthetically qualified the title of this post. References to "primary texts" - and other texts for that matter - are not simple. In this post - as is often my wont - I've let myself be drawn along by current practice. I really do like to see what people are actually doing and how data actually works on the Internet. If you want a more substantive discussion of the problems of citation, I highly recommend Neel Smith's "Citation in Classical Studies" in DHQ 2009. Here's the abstract:
Citation practice reflects a model of a scholarly domain. This paper first considers traditional citation practice in the humanities as a description of our subjects of study. It then describes work at the Center for Hellenic Studies on an architecture for digital scholarship that is explicitly based on this model, and proposes a machine-actionable but technologically independent notation for citing texts, the Canonical Text Services URN.
For now, let me say that it is correct for Google (via Biblegateway) to dereference a citation to John 7:53-8:11 (the Pericope Adulterae) or John 5:7 (the Comma Johanneum). Neither may have been in the "original" text of the Gospel of John, but references to them are semantically clear and have been used "in the wild" so need to be handled. But note that Google prioritizes discussion of the CJ over the text (or at least does when I'm trying it now). Again, see N. Smith on the implications of such variation.

Clearly it helps to have a committed body of believers and/or scholars working on very old texts. Energy and time make for stable references. But there is variability in functionality even within that group. I guess the long-term question is how do we move more texts into the category of "just working"? I am assuming we want to. And how do we support co-existence of the simple "reference following" alongside what Neel describes. Both are useful.

Monday, August 30, 2010

Numbered Paragraphs in Digital Humanities Quarterly

I can recommend Patrik Svensson's article "The Landscape of Digital Humanities" in Digital Humanities Quarterly as a good read. My comments here are about the internals of handing DHQ's paragraph based citation scheme.

Quick intro to the issue: DHQ is an online journal. It doesn't have pages to provide a physical solution to the need to make references to specific points in an article. So the html version numbers each paragraph. So far so good. As a reader I can note the paragraph number and cite it in a future publication.

But I'm not sure DHQ has quite the right implementation of this good idea. I'm arbitrarily picking the paragraph numbered 118. The one that starts, "Information technology, or more broadly the digital, can be seen as affording objects of analysis for the humanities."

Note that I don't include a link directly to that paragraph. That's because I can't. Looking at the HTML source, I see:
<div class="counter">118</div><div class="ptext">Information technology, or more broadly the digital, can be seen as affording...


That's somewhat unfortunate. It would be great if the '<div class="ptext">' were changed to read '<div id="p118" class="ptext">. Then I could mint a URL of the form:
http://www.digitalhumanities.org/dhq/vol/4/1/000080/000080.html#p118


It would be even cooler if the <div class="counter">118</div> also read:
<div class="counter"><a href="#p118">118</a></div>


I've wrapped the paragraph number in a link to the paragraph. That way a user can right/control-click on the link and copy-and-paste it into an e-mail or other work. Easy self-reference to an internal citation structure.

I'd also like to see the paragraph numbers represented in the XML source. Again, taking a snippet of that, the start of the paragraph numbered as 118 in the html, appears in the xml as:
<p>Information technology, or more broadly the digital, can be seen as affording objects of analysis for the humanities...


Unless I'm missing something, the published citation scheme isn't represented in the archival version. I think it should be. Even if DHQ considers the paragraph number ephemeral, I think there's a valid scholarly need for them to be persistent.

I'm a big fan of DHQ so this is constructive criticism. And I'm sort of hoping that I've mis-understood something and that those paragraph numbers are more meaningful than they seem after one looks under the hood.

Tuesday, August 24, 2010

Corrected Versions of Papers

It's late August so my mind is on other things, like the next stretch of split-rail fence that I need to put in. But I do find it interesting that Heather Baker has used Academia.edu to distribute a corrected version of her paper "The layout of the ziggurat temple at Babylon" that first appeared in Nouvelles Assyriologiques Brèves et Utilitaires 2008.2 (Juin). Feel free to be similarly and vaguely inspired about issues of versioning, "scribal error", reference, etc.

Thursday, June 17, 2010

More Papers on Academia.edu

Sure, Academia.edu is far from perfect. But I continue to be psyched when I see people who have uploaded a bunch of papers or a book.Just a brief "thanks" to those who have added to my collection of digital offprints.

Thursday, June 10, 2010

Academia.edu

When I first heard about Academia.edu, I pretty much ignored it. Basically, the last thing I needed was another social media site.

A few weeks ago, however, Chuck Jones sent me an invitation so I took another look. What caught my eye was the papers that users have uploaded. I'm always on the look out for digital content that I can't find elsewhere so the site's role as a central point for the discovery of articles, etc. is very useful.

Here's an incomplete list of some of my fellow "Academicians" whose scholarship I've downloaded:I've tried to do my bit as well by pointing to some of my work that's available on-line. And here are links to the Institute for the Study of the Ancient World and American Numismatic Society pages. I'm intrigued by the opportunity to have a lightweight institutional repository that Academia.edu offers to an organization like the ANS.

Not everything is perfect about the site. The 'Department Viewer' - for want of a better name - relies on Flash. That's lame and doesn't work on an iPad. And I'm not sure the developers have realized that the papers are a real asset. It would be nice to be able to see all papers listed by people I'm following. Or all papers from a single department. And more liberal spreading around of "Follow this Person" buttons would be nice. If I'm looking at the page for a particular paper, I should be able to single-click to follow its author.

If you're interested, come join the fun. And remember to link to your digital scholarship. As you can tell, I think that's the real utility of the site.

Coin Hoards, Timelines and KML

http://nomisma.org/kml/thasos-all.kml is a KML file that shows findspots of hoards with coins of Thasos in them. You can see that file rendered with the Google Map API at http://nomisma.org/id/thasos.

This post is about viewing the KML file in Google Earth. If you do that, you'll see a Timeline slider appear in the top left of the G Earth window. Slide the control to the right and you'll see an explosion of hoards towards the north following the mid-2nd century BC. It's really quite dramatic so give it a shot.

One part of an explanation is that Thasos started striking large numbers of new larger tetradrachms following 148 BC, with many of these traveling north. Exactly why is a matter of historical interpretation. The Roman province of Macedon was established in 146 B.C. and that had a profound effect on both the issuance and circulation of coinage.

Nomisma.org is about making this information more accessible so that more scholars can engage with the question.

Wednesday, June 2, 2010

Bibliographic Tools, Citations, and Digital Publications

Some preliminaries... I'm posting about digital publication of ancient world scholarship as part of my work at NYU's Institute for the Study of the Ancient World. I say that only to make clear that there's a practical element to my thinking about issues of citation, structure, metadata, etc. I will be helping to shepherd content into the digital realm and that means decisions, decisions, decisions. I am enjoying the focus this context gives to my thoughts. And since I, like most bloggers, live for little nuggets of feedback, those have been appreciated as well.

It's also important to stress that this is all happening within the intellectual context of ISAW. In other words, my new colleagues have pushed on these issues in interesting ways and I can take advantage of their previous work.

As in, let's talk about Zotero and the role it can play in providing a sustainable bibliographic framework for digital scholarship. This is already happening at the ISAW-hosted Pleiades project, and that lets me take a very practical approach to writing about the issue.

The Pleiades Zotero library is at http://www.zotero.org/groups/pleiades/items. It includes the item http://www.zotero.org/groups/pleiades/items/72922905, which is the Zotero record for the article
Coastal Sites of Northeast Africa: The Case Against Bronze Age Ports
. In case it isn't clear, the point of the article is pretty much that there weren't any.

In yesterday's post, I talked about citing secondary scholarship. Today, I'm interested in the mediating those citations through Zotero bibliographic records.

The same basic pattern would apply: '<a rel="dc:references" href="http://www.zotero.org/groups/pleiades/items/72922905">White and White 1996</a>' is a reference to the work described in that Zotero record. I am interested in the extent to which it is necessary to indicate that it is not a reference to the Zotero record itself but let me put that off for now. More relevant here is why use Zotero to establish unique identities for cited works at all?

The most compelling reason is that not all works will have such identifiers and Zotero allows one to create these. For example, http://www.zotero.org/groups/pleiades/items/72922931 is the Zotero record for an article that has no direct representation in Worldcat and which isn't online (I don't think). I.e., you're on your own in terms of a stable URI for this title.

And since consistency is good, it might be appropriate to create Zotero records for all cited works in a digital publication and only point to those.

This approach is also attractive because it allows linking to digital representations of titles as they become available. For example, the record for Coastal Sites could be linked to http://www.jstor.org/stable/40000602, which is the JSTOR record. A more compelling example is this link: http://hdl.handle.net/10.2972/hesp.79.1.1. That will take you to the Atypon-Link version of J. Cherry and W. Parkinson's Hesperia 2010 article on lithics from SW Greece. As the volumes of Hesperia role over into JSTOR once they are past the 3 year wall, the Atypon URI will be either matched or replaced by an equivalent JSTOR URI. A Zotero record can have links to both versions without requiring any updating of the digital publication which points to that record.

And if more than one digital publication points to a Zotero record, that equivalency should be discoverable. I like that.

The big potential downer is: do we trust Zotero to be around for the foreseeable future? Or at least, will these URI's work over the very long term? I don't know the answer to that. This is one reason to ensure that each digital publication "knows" bibliographic metadata for all the citations it makes. Centuries from now, that information may be useful in tracking down readable versions of titles.

And here'a a finishing twist. Regardless of which tool is used to generate URI-based unique identifiers for cited works, that same tool could (should? must?) be used to provide URI-based unique identifiers for the digital publication itself.

Tuesday, June 1, 2010

References in Digital Publications

Modern scholarship relies on citation. It's efficient in that one work can incorporate the results of another without having to repeat it. It's also a requirement of our modern academic culture that if you use somebody's idea, you give that person credit. There's more to be said on both points but this post is more about mechanics than purpose. (Though see here for a recent discussion of purpose. [I fall into the camp of : if you want credit for your work, make it easy to identify and be generous in giving credit to others. If you don't need credit, that's OK but still give it.]).

Back to references. They come in many forms in print works. In pre-linked media, among the purposes of citation is to give future readers the information they need to physically acquire the referenced work. That is, you take the title of the book or journal, go to the library to find the volume, and then start reading.

It is one of the great glories of the Internet that this physical labor is no longer always necessary. The simple construct '<a href="http://sebastianheath.com/files/HeathS2010-DigitalResearch.pdf">I wrote this</a>' is rendered as 'I wrote this', so that a mere click takes you directly to the article.

That form of link is too simple to support modern scholarly practice. Citations of the form (Heath 2010) give a preliminary indication to the reader of who wrote a referenced work. Full information in footnotes further enriches the reading experience, but at the cost of possibly interrupting the flow of an argument, or depriving the reader of a collected bibliography at the end of a work. Choose your own preference, that's not my point here.

Instead, I am exploring specific patterns of markup that promote access to referenced works while also recording bibliographic metadata in a robust and sustainable fashion. Two needs, two solutions.

Here's some markup: Late Roman pottery is very visible in Aegean landscapes (<a rel="dcterms:references" href="http://hdl.handle.net/10.2972/hesp.76.4.743">Pettegrew 2007</a>).

If we momentarily ignore the question of whether or not Handles records are good stable URIs for bibliographic resources, the semantics of this html are clear: it represents a citation of the 2007 article by David Pettegrew, The Busy Countryside of Late Roman Corinth. (Note: it doesn't reference the html page describing that title)

The use of the term "dcterms:references" in the RDFa rel attribute follows from the Dublin Core's Guidelines for Encoding Bibliographic Citation Information in Dublin Core Metadata. In this context 'references' is a verb, not a plural noun.

That html will render as: "Late Roman pottery is very visible in Aegean landscapes (Pettegrew 2007)." Again, this is all pretty clear.

It's also worth noting that the 'a' element in html is a building-block of our search-engine enabled world. Scholarship should not fight that, but use it. As many have said, "you get this for free."

I do, however, want to pair this reference with bibliographic metadata. Here's where some more RDFa comes in.

'http://hdl.handle.net/10.2972/hesp.76.4.743' is a unique identifier for Pettegrew's article. This suggests the following snippet: <div about="http://hdl.handle.net/10.2972/hesp.76.4.743"><span property="dcterms:bibliographicCitation">Pettegrew, D. (2007). "The Busy Countryside of Late Roman Corinth: Interpreting Ceramic Data Produced by Regional Archaeological Surveys" In <i>Hesperia</i> 76.4: 743-784.</span></div>

These two snippets can be adapted and combined with a little more RDFa scaffolding:
<html xmlns="http://www.w3.org/1999/xhtml"
xmlns:dcterms="http://purl.org/dc/terms/" >
<body about="http://example.org/example_document">
<h1>My Text</h1>
<p>Late Roman pottery is very visible in Aegean landscapes (<a rel="dcterms:references" href="http://hdl.handle.net/10.2972/hesp.76.4.743">Pettegrew 2007</a></p>
<h1>References</h1>
<p about="http://hdl.handle.net/10.2972/hesp.76.4.743" property="dcterms:bibliographicCitation">Pettegrew, D. (2007). "The Busy Countryside of Late Roman Corinth: Interpreting Ceramic Data Produced by Regional Archaeological Surveys" In <i>Hesperia</i> 76.4: 743-784.</p>
</body>
</html>


Pointing an RDFa extractor at that html gives:
@prefix rdf: <http://www.w3.org/1999/02/22-rdf-syntax-ns#> .
@prefix : <http://www.w3.org/1999/xhtml> .
@prefix dcterms: <http://purl.org/dc/terms/> .

<http://example.org/example_document>
   dcterms:references <http://hdl.handle.net/10.2972/hesp.76.4.743> .

<http://hdl.handle.net/10.2972/hesp.76.4.743>
   dcterms:bibliographicCitation "Pettegrew, D. (2007). \"The Busy Countryside of Late Roman Corinth: Interpreting Ceramic Data Produced by Regional Archaeological Surveys\" In <i xmlns=\"http://www.w3.org/1999/xhtml\" xmlns:dcterms=\"http://purl.org/dc/terms/\">Hesperia</i> 76.4: 743-784."^^rdf:XMLLiteral .


The shorter version of which is: example.org/example_document references Pettegrew 2007 and even knows something about it. There are lots of third-party tools that can find this information when it is encoded in this way. And I could enrich the 'bibliographicCitation' to include parsable information on author, title, date, etc. That's for another time.

I want to stress that I don't think this determines a particular citation style. Use footnotes if that's preferable. As long as the RDFa produces triples similar to the above, your information is useful. And some degree of run-time transformation is also possible, depending on the granularity of the markup.

Tuesday, May 25, 2010

Me @ NYU/ISAW

Briefly... I'm sitting in an office at NYU's Institute for the Study of the Ancient World, where I am now a Visiting Scholar. This is preliminary to a more permanent position with details to come.

My main goal is to work on issues of digital publication and on integration of diverse digital resources. I had started collaborating with ISAW-folk on these issues some time back, which is why I've been blogging about them.

I'm extremely excited to be working with my new colleagues here - a veritable dream-team of digital humanists - and am looking forward to making real progress when it comes to sharing well-structured, semantically-rich, open-licensed scholarship about the Ancient World.

And I'll still be collaborating with my long-time friends at the ANS, particularly on Nomisma.org. And field-work goes on.

Back to work...

Wednesday, May 19, 2010

RDFa Document Metadata: Authors in PLOS One

Brief follow up to yesterday's post.

Here's the HTML that indicates authorship from an example PLOS One article.
<p xmlns:xs="http://www.w3.org/2001/XMLSchema" xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:aml="http://topazproject.org/aml/" class="authors" xpathlocation="noSelect">
<span rel="dc:creator"><span property="foaf:name">Harold C. Sox</span></span><sup><a href="#aff1">1</a></sup>, <span rel="dc:creator"><span property="foaf:name">Mark Helfand</span></span><sup><a href="#aff2">2</a></sup><sup><a href="#cor1" class="fnoteref">*</a></sup>,
<span rel="dc:creator"><span property="foaf:name">Jeremy Grimshaw</span></span><sup><a href="#aff3">3</a></sup>,
<span rel="dc:creator"><span property="foaf:name">Kay Dickersin</span></span><sup><a href="#aff4">4</a></sup>, <span class="capture-id">the <i>PLoS Medicine</i> Editors</span>,
<span rel="dc:creator"><span property="foaf:name">David Tovey</span></span><sup><a href="#aff5">5</a></sup>, <span rel="dc:creator"><span property="foaf:name">J. André Knottnerus</span></span><sup><a href="#aff6">6</a></sup>,
<span rel="dc:creator"><span property="foaf:name">Peter Tugwell</span></span><sup><a href="#aff7">7</a></sup>
</p>
<p xmlns:xs="http://www.w3.org/2001/XMLSchema" xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:aml="http://topazproject.org/aml/" class="affiliations" xpathlocation="noSelect">
<a name="aff1" id="aff1"></a><strong>1</strong> Dartmouth Institute, Dartmouth Medical School, Hanover, New Hampshire, United States of America,
<a name="aff2" id="aff2"></a><strong>2</strong> Portland VA Medical Center and Department of Medicine, Oregon Health &amp; Science University, Portland, Oregon, United States of America,
<a name="aff3" id="aff3"></a><strong>3</strong> Clinical Epidemiology Program, Ottawa Hospital Research Institute, Ottawa, Ontario, Canada,
<a name="aff4" id="aff4"></a><strong>4</strong> Department of Epidemiology, Johns Hopkins Bloomberg School of Public Health, Baltimore, Maryland, United States of America,
<a name="aff5" id="aff5"></a><strong>5</strong> The Cochrane Library, London, United Kingdom, <a name="aff6" id="aff6"></a>
<strong>6</strong> Department of General Practice, University of Maastricht, Maastricht, The Netherlands,
<a name="aff7" id="aff7"></a><strong>7</strong> Departments of Medicine, and Epidemiology and Community Medicine, University of Ottawa, Ottawa, Ontario, Canada
</p>


The basic structure is two 'p' elements, one with a 'class="authors"', the second with 'class="affiliations"'. I am trying to avoid using @class to indicate document structure and metadata, so yesterday I adopted the 'bibo:authorList' convention. But it is useful to see another instance of the nested 'rel="dc:creator"'->'property="foaf:*"' pattern. Is that beginning to look like a trend?

The relationship between author and affiliation is a little broken. The reference from each author to his/her affiliation is actually to an 'a' element with no content. An automatic agent might return an empty string as the affiliation unless it had ad hoc code to pull the text as far as the next '<a>' or '</p>' tag. That's not particularly helpful.

It is important to be clear that this HTML is rendered from XML encoded in the National Institutes of Health's Journal Publishing Tag Set Version 2.0. That's my way of acknowledging that the markup delivered to your browser doesn't bear the full weight of being a well-structured archival version.